====================================================================================
MerQur - Anestezi_Reanimasyon - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Anestezi_Reanimasyon/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_patient_monitor.xlsx
  >> SCENARIO (narration):
    We compiled the demographic and monitoring inventory of 280 patients. Age, height,
    weight, BMI, arterial pressure and glucose were measured for each. Before any
    inferential test we want the overall picture of the sample; so we begin with
    descriptive statistics.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: height_cm
    - Variables: weight_kg
    - Variables: BMI
    - Variables: MAP_mmHg
    - Variables: CVP_cmH2O
    - Variables: glucose_mg_dL
    - Grouping (categorical): sex

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 patients
    age: mean = 51.9   |   height: mean = 167.8 cm   |   weight: mean = 76.1 kg   |   BMI: mean = 27.3
    mean arterial pressure: mean = 134.0 mmHg   |   glucose: mean = 96.1 mg/dL

>> COMMENTARY (narration):
    First we draw the overall picture of our patient sample: 280 patients, average age ~52, height 168 cm, weight 76 kg,
    BMI 27.3 (borderline overweight), mean arterial pressure 134 mmHg (mildly elevated), glucose 96 mg/dL. This descriptive table
    lays the groundwork for every analysis that follows -- group/anesthetic comparisons, clinical-lab relationships, spatial
    disease pattern. In anesthesia and critical care, before any inferential test, summarizing the patient sample's basic features (age,
    hemodynamics, labs) is essential both to audit data quality and to set clinical priorities.

====================================================================================

#2  Normality Tests
    file: 02_normality_BMI_MAP_SpO2.xlsx
  >> SCENARIO (narration):
    We examine whether BMI, mean arterial pressure and SpO2 are normally distributed. Because
    subsequent t-tests, ANOVA and correlation depend on this assumption, we test
    each variable separately.
  >> VARIABLE SELECTION:
    - Variables: BMI
    - Variables: MAP_mmHg
    - Variables: SpO2_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    3 continuous variables tested:
    BMI               : Shapiro-Wilk = 0.996  p = 0.881   KS p = 0.917   Normal
    MAP (mmHg)  : Shapiro-Wilk = 0.991  p = 0.275   KS p = 0.252   Normal
    SpO2_pct         : Shapiro-Wilk = 0.939  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three continuous variables -- BMI, mean arterial pressure and SpO2 -- are normally distributed.
    The result splits instructively: BMI and MAP are normal (p > 0.27), but SpO2 deviates significantly from normality
    (p < .001). This is typical in clinical data -- SpO2 is often right-skewed (most patients normal, a few critically ill patients
    form a high tail). Practical upshot: we can safely use parametric tests (t-test, ANOVA, Pearson) on BMI and MAP; for
    SpO2, nonparametric methods (Mann-Whitney/Kruskal-Wallis) or a transform are more appropriate. The normality check
    is a critical preliminary step that decides, per variable, which test family fits.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_MAP_bp.xlsx
  >> SCENARIO (narration):
    We investigate whether the patients' mean mean arterial pressure differs from the 120 mmHg
    normal reference. With one group and a fixed reference, the one-sample t-test is
    appropriate.
  >> VARIABLE SELECTION:
    - Test variable: MAP_mmHg
    - Test value (mu): 120

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = 5.860   p < .001 ***   Cohen d = 0.586 (Medium)
    Mean mean arterial pressure = 128.25 mmHg   (test mu = 120, normal reference)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the patients' mean mean arterial pressure against the normal reference of 120 mmHg. The result is
    significant and medium-sized: mean 128.25 mmHg, above the normal limit -- t(99) = 5.86, p < .001, d = 0.59. So this
    patient group is above normal on average, with a prehypertensive tendency. The one-sample t-test is the right way to
    compare a clinical measure against a known reference/threshold (normal MAP, target value); in medicine it is widely
    used to evaluate a group against a clinical norm.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent-Samples t-Test
    file: 04_independent_t_anesthetic_SpO2.xlsx
  >> SCENARIO (narration):
    We compare the mean SpO2 change of two anesthetic groups. With two separate groups
    and a continuous measure, the independent-samples t-test is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): group
    - Test variable: SpO2_change

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = -10.82   p < .001 ***   Cohen d = -2.019 (Large)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the mean SpO2 change of two anesthetic groups. The difference is significant and very large: t(113) = -10.82,
    p < .001, d = -2.02. It shows the between-group difference is too pronounced to be chance and is also clinically
    noteworthy. The independent-samples t-test is the standard way to compare the means of two separate groups
    (anesthetic/placebo, two arms) on a continuous measure; in clinical research it is a fundamental tool for detecting
    differences between treatment arms.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired-Samples t-Test
    file: 05_paired_t_MAP_treatment.xlsx
  >> SCENARIO (narration):
    We compare mean arterial pressure measured before and after treatment in the same patients.
    Since the measures are paired, the paired t-test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: MAP_before_mmHg
    - 2nd measure / group: MAP_post_mmHg

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(49) = 18.43   p < .001 ***   Cohen d_z = 2.607 (Large)
    Mean diff (before - after) = 19.98 mmHg   H0 REJECTED

>> COMMENTARY (narration):
    We paired and compared mean arterial pressure measured before and after treatment in the same patients. The result
    is very strong: mean difference 19.98 mmHg (a drop), t(49) = 18.43, p < .001, d_z = 2.61, a huge effect. So MAP
    dropped significantly and substantially after treatment. The paired t-test compares two timed measures on the same
    patient (before/after treatment); by isolating individual change it is more powerful than the independent test and
    is the right way to measure a treatment effect in medicine.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_anesthetic_EtCO2.xlsx
  >> SCENARIO (narration):
    We compare the effect of four anesthetic groups on mean EtCO2 change. With more than two
    groups, one-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: EtCO2_change
    - Factor (categorical): anesthesia_technique

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,136) = 57.67   p < .001 ***   eta^2 = 0.560   H0 REJECTED

>> COMMENTARY (narration):
    We compared mean EtCO2 change across four anesthetic groups. The result is significant and very strong: F(3,136) = 57.67,
    p < .001, eta^2 = 0.56 -- half of EtCO2-change variance comes from anesthetic differences. At least one anesthetic differs
    significantly. One-way ANOVA compares the means of more than two groups at once (avoiding the error inflation of
    many t-tests); in medicine it is the core method for comparing different anesthetic/dose/arm effects. Which pairs differ
    is then determined by post-hoc tests.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova_anesthetic_sex.xlsx
  >> SCENARIO (narration):
    We examine the main effects and interaction of anesthetic and sex on blood-pressure
    change simultaneously. With two categorical factors, two-way ANOVA is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: bp_change
    - Factor (categorical): anesthetic
    - 2nd Factor: sex

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    anesthetic (main effect): F(2) = 29.95   p < .001 ***   eta^2p = 0.294
    sex (main effect) : F(1) = 9.04    p = 0.003 **    eta^2p = 0.059

>> COMMENTARY (narration):
    We examined two factors at once: how do anesthetic and sex affect blood-pressure change? The anesthetic main effect is very
    strong (F(2) = 29.95, p < .001, eta^2p = 0.29), and the sex main effect is also significant but small (F(1) = 9.04,
    p = 0.003, eta^2p = 0.06). So the dominant driver of MAP change is anesthetic; sex makes a modest contribution. The power
    of two-way ANOVA is that it tests both factors and their interaction in a single model -- isolating each factor's
    pure effect with the other controlled. It is ideal for answering "which variable really makes a difference?" in
    medicine.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated-Measures ANOVA
    file: 08_repeated_anova_visit_SpO2.xlsx
  >> SCENARIO (narration):
    We compare SpO2 over four consecutive visits in the same patients. With
    repeated measures on the same patient, repeated-measures ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: visit_1_SpO2_pct
    - Repeated measures: visit_2_SpO2_pct
    - Repeated measures: visit_3_SpO2_pct
    - Repeated measures: visit_4_SpO2_pct

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,237) = 155.68   p < .001 ***   eta^2p = 0.663   n = 80   H0 REJECTED

>> COMMENTARY (narration):
    We compared SpO2 measured across four consecutive visits in the same 80 patients. The result is very strong:
    F(3,237) = 155.68, p < .001, eta^2p = 0.66 -- the between-visit difference is huge and most of the effect is
    time-related. So patients' SpO2 changes significantly across follow-up. Repeated-measures ANOVA compares three or
    more timed measurements on the same patient; by holding individual differences constant it yields high statistical
    power and is the right choice for multi-visit follow-up (SpO2 tracking, periodic monitoring) in medicine.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_treatment_3monitor.xlsx
  >> SCENARIO (narration):
    We test treatment's effect on MAP, CVP and heart rate at once. With
    several correlated dependent variables, MANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: MAP_mmHg
    - Dependent variable: CVP_cmH2O
    - Dependent variable: heart_rate_bpm
    - Factor (categorical): treatment

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.346   F(6,230) = 26.83   p < .001 ***   n = 120
    Dependent: MAP, CVP, heart rate   |   Factor: treatment   H0 REJECTED

>> COMMENTARY (narration):
    We tested treatment's effect on three hemodynamic monitoring (MAP, CVP, heart rate) simultaneously. With Wilks' Lambda =
    0.35, F(6,230) = 26.83, p < .001, the effect is very strong. MANOVA examines several correlated outcomes in one
    test, both preventing the error inflation of many separate ANOVAs and capturing the joint information the variables
    carry together. In anesthesia and critical care, when treatment affects not a single hemodynamic but the hemodynamic "bundle" (mean arterial pressure-central venous pressure-
    heart rate) together, MANOVA reveals this multivariate difference as a single decision.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova_physical_treatment_VAS.xlsx
  >> SCENARIO (narration):
    We compare post-treatment pain across groups while controlling baseline pain as
    a covariate. With a confounding continuous variable, ANCOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: pain_post_VAS
    - Factor (categorical): treatment
    - Covariate: pain_baseline_VAS

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group effect significant   eta^2p = 0.712 (very large)   covariate: pain_baseline_VAS   H0 REJECTED

>> COMMENTARY (narration):
    We compared post-treatment pain (pain_post_VAS) across groups while controlling baseline pain (pain_baseline_VAS) as
    a covariate. This rules out the objection that "the groups differed at baseline" and measures the pure group effect:
    the effect is very large (eta^2p = 0.71). ANCOVA makes the group comparison fair by statistically holding a
    confounding continuous variable constant -- it is the answer to "once we equalize baseline pain, does the treatment
    still make a difference?" and is the standard tool in baseline/post-measure designs in medicine.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci_lactate.xlsx
  >> SCENARIO (narration):
    For mean blood lactate we build a confidence interval via resampling, with no
    distributional assumption. For skewed data, bootstrap is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: lactate_mmol_L

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean (lactate, mmol/L) = 2.06
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    For the mean blood lactate (mmol/L) we produced a 95% confidence interval via resampling -- with no distributional
    assumption: observed mean 2.06 mmol/L. Bootstrap builds the sampling distribution of the statistic empirically by
    resampling the data thousands of times from itself; it is a reliable way to give a confidence interval when
    normality does not hold (lactate is skewed) or no formula is known. In medicine it provides more robust estimates
    than classic t-intervals for skewed laboratory measures.

====================================================================================

#12  Permutation Test
    file: 12_permutation_procalcitonin_case_control.xlsx
  >> SCENARIO (narration):
    We test the CRP difference between two groups (case/control) via permutation,
    with no distributional assumption. For small samples/skewed distributions,
    permutation is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: procalcitonin_ng_mL
    - Grouping (categorical): group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = -3.472   p = 0.001 **   significant mean difference between the two groups   H0 REJECTED

>> COMMENTARY (narration):
    We tested the CRP difference between two groups (case/control) without any distributional assumption, via
    permutation: the null distribution was built by randomly swapping group labels, and the observed difference turned
    out rare in that distribution (p = 0.001). Because the permutation test is exact and distribution-free, it is a safe
    alternative to the parametric t-test for small samples or skewed laboratory distributions; in medicine it is a
    robust choice for case-control comparisons.

====================================================================================

#13  Multiple Comparison
    file: 13_multiple_comparison_antikoagulan.xlsx
  >> SCENARIO (narration):
    We take the p-values of eight anticoagulant comparisons together and apply
    multiple-testing correction. With many tests, p-adjustment is appropriate.
  >> VARIABLE SELECTION:
    - Variables: anesthetic
    - Variables: INR_stability_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Raw: 4/8 significant   |   After Bonferroni / Holm / BH (FDR): 4/8 significant (all retained)

>> COMMENTARY (narration):
    We took the p-values of eight separate anticoagulant/anesthetic comparisons together and applied multiple-testing
    correction. Raw, 4 comparisons were significant; after Bonferroni/Holm/BH all 4 remained significant -- so these
    differences are strong enough to survive correction. When many tests are run, the rate of false positives that look
    "significant" by chance alone inflates; correction methods tighten the threshold to control this error. In anesthesia and critical care,
    when many anesthetic agents/doses are compared at once, this step protects inference reliability.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney_diabetes_survival_quality.xlsx
  >> SCENARIO (narration):
    We compare the survival-quality distribution of two diabetes types. Since
    normality fails, Mann-Whitney is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): diabetes_type
    - Test variable: survival_quality

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    U = 886.50   p = 0.907 ns   r = 0.02   H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the survival-quality distribution of two diabetes types (diabetes_type) based on ranks rather than
    means: U = 886.5, p = 0.907, a negligible effect (r = 0.02) -- no significant difference. Mann-Whitney is the
    nonparametric counterpart of the t-test; when normality fails or the scale is ordinal, it is the right way to
    compare two groups. Here the "no difference" result tells us the two diabetes types are similar in survival quality.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank
    file: 15_wilcoxon_treatment_VAS.xlsx
  >> SCENARIO (narration):
    We compare pain scores measured before/after treatment in the same patients,
    without assuming normality. For paired non-normal data, Wilcoxon is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: VAS_before
    - 2nd measure / group: VAS_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    W = 0.00   p < .001 ***   r = 0.89   n (non-zero) = 31   H0 REJECTED

>> COMMENTARY (narration):
    We compared pain scores measured before and after treatment in the same patients (VAS_before vs post) -- without
    assuming normality, on a rank basis: W = 0, p < .001, very large effect (r = 0.89). So pain changed consistently and
    strongly after treatment. Wilcoxon is the nonparametric counterpart of the paired t-test; it is the right choice for
    ordinal or non-normal before/after measures. In medicine it gives reliable results for pre/post-treatment VAS pain
    measures.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_vasopressor_duration.xlsx
  >> SCENARIO (narration):
    We compare the treatment-duration distribution across three antibiotics. Since
    normality/variance homogeneity fails, Kruskal-Wallis is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): antibiotic
    - Test variable: treatment_day

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 3.35   p = 0.187 ns   eta^2_H = 0.03   H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the treatment-duration (treatment_day) distribution across three antibiotics by ranks rather than means:
    H(2) = 3.35, p = 0.187, eta^2_H = 0.03 -- no significant difference. Kruskal-Wallis is the nonparametric counterpart
    of one-way ANOVA; when normality or variance homogeneity fails, it is the right way to do multi-group comparison.
    Here the "no difference" result tells us the antibiotics are similar in treatment duration.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman_radiologist_method.xlsx
  >> SCENARIO (narration):
    We compare the accuracy of four imaging methods in the same cases. As a
    nonparametric repeated measure, Friedman is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: method_A_accuracy
    - Repeated measures: method_B_accuracy
    - Repeated measures: method_C_accuracy
    - Repeated measures: method_D_accuracy

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 32.08   p < .001 ***   Kendall W = 0.356   n = 30   H0 REJECTED

>> COMMENTARY (narration):
    We compared the accuracy of four imaging methods (method_A..D) in the same 30 cases -- as a nonparametric repeated
    measure: chi^2(3) = 32.08, p < .001, Kendall W = 0.36 (moderate concordance). The between-method difference is
    significant. Friedman is the nonparametric counterpart of repeated-measures ANOVA; it is the right choice for
    ordinal or non-normal paired multi-condition measures. In medicine it is used to compare the performance of
    different diagnostic methods on the same patient/case.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_surgery_achievement.xlsx
  >> SCENARIO (narration):
    We test the proportion of successful surgeries against an expected 50%. With a
    binary outcome and a theoretical proportion, the binomial test is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: successful
    - Expected proportion: 0.50
    - Success value: 1

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed proportion = 0.738   (expected p = 0.50)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested the proportion of successful surgeries against an expected 50%: observed proportion 73.8%, well above
    expectation (p < .001). So the surgical success rate is too high to be chance. The binomial test is the exact method
    for comparing the observed proportion of a binary (success/failure) outcome with a theoretical proportion; in
    medicine it directly tests whether procedure/treatment success rates meet a target or a 50:50 expectation.

====================================================================================

#19  Sign Test
    file: 19_sign_test_diet_weight.xlsx
  >> SCENARIO (narration):
    We look at the direction of before/after weight on a diet in the same patients.
    When only directional information is reliable, the sign test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: weight_before_kg
    - 2nd measure / group: weight_post_kg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive diffs (before > after) = 26   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We looked at the direction of change in weight measured before/after a diet in the same patients (weight_before vs
    post): the large majority of changes go in one direction (a decrease), p < .001 for a significant directional
    change. The sign test uses only the direction of the difference (increase/decrease), not its magnitude, so it is the
    before/after test that requires the fewest assumptions. In medicine it is a robust choice when the measure is skewed
    or only directional information is reliable (did weight drop or rise).

====================================================================================

#20  Runs Test
    file: 20_runs_test_YBU_occupancy.xlsx
  >> SCENARIO (narration):
    We test whether the bed-occupancy sequence (around the median) is random. For
    sequence randomness, the runs test is appropriate.
  >> VARIABLE SELECTION:
    - Column: bed_occupancy

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 24   Z = -1.816   p = 0.069 ns   data may be considered random

>> COMMENTARY (narration):
    We tested whether the binary sequence of bed occupancy (bed_occupancy, around the median) is random: 24 runs,
    Z = -1.82, p = 0.069 -- borderline but no significant pattern, the sequence can be considered random. The runs test
    checks whether values in a sequence form a systematic pattern (clusters, cycles, trend). In medicine it is used to
    determine whether systematic patterns or randomness dominate in time series like bed occupancy and case series.

====================================================================================

#21  Chi-Square Independence
    file: 21_chisquare_smoking_copd.xlsx
  >> SCENARIO (narration):
    We test whether smoking status and COPD are related. With two categorical
    variables, chi-square is appropriate.
  >> VARIABLE SELECTION:
    - Variables: smoking_status
    - Variables: copd

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(2) = 130.54   p < .001 ***   Cramer's V = 0.441 (Strong)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether two categorical variables -- smoking status (smoking_status) and COPD (copd) -- are related:
    chi^2(2) = 130.54, p < .001, Cramer's V = 0.44, a strong dependency. The categories are not independent; they vary
    together. The chi-square test of independence detects the relationship between two qualitative variables from a
    cross-tab; in medicine it is the core method for revealing risk factor-disease (smoking-COPD) categorical
    relationships. Cramer's V measures the practical strength of the relationship.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_blood_group.xlsx
  >> SCENARIO (narration):
    We test whether a four-category blood group's observed distribution fits an
    equal expected distribution. For one categorical variable, goodness-of-fit is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: blood_group

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 88.12   p < .001 ***   N = 200, k = 4 (expected: equal distribution)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the observed distribution of a four-category variable (blood group) fits an equal (1/k) expected
    distribution: chi^2(3) = 88.12, p < .001 -- the blood groups are not equally distributed, some are clearly more
    frequent. The goodness-of-fit test compares the observed frequencies of a single categorical variable with a
    theoretical expectation (equal proportions, a known population ratio); in medicine it tests whether blood-group/
    diagnosis distributions match an expected profile.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_diagnosis_test.xlsx
  >> SCENARIO (narration):
    In a 2x2 table we test the test-result relationship exactly, due to small cell
    frequencies. For few observations, Fisher is appropriate.
  >> VARIABLE SELECTION:
    - Variables: test
    - Variables: result

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher exact p = 0.276 ns   Odds Ratio = 0.389   Phi = 0.227 (Moderate)   H0 NOT REJECTED

>> COMMENTARY (narration):
    In a 2x2 cross-tab (test x result) we tested the relationship exactly -- using Fisher instead of chi-square because
    of small cell frequencies: p = 0.276, no significant relationship (Phi = 0.23). Fisher's exact test is the right
    choice when expected frequencies are low and the chi-square approximation is unreliable; it computes the probability
    exactly rather than approximately. In medicine it gives reliable results for small-sample diagnosis-test
    comparisons; here test and result turned out independent.

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_old_new_diagnosis.xlsx
  >> SCENARIO (narration):
    We compare the binary outcome of old/new diagnostic methods paired in the same
    patients. For paired binary change, McNemar is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: old_method
    - 2nd measure / group: new_method

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar (exact) min(b,c) = 4   p = 0.077 ns   discordant b = 12, c = 4   OR = 3.00   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested the direction of change in a binary outcome of two diagnostic methods (old_method vs new_method) paired in
    the same patients: although the changes are imbalanced (b = 12, c = 4), p = 0.077 is borderline but non-significant.
    McNemar tests the old/new diagnostic-result change in the same patient and looks only at discordant pairs. In
    medicine it is the right tool for detecting whether a new diagnostic method gives systematically different results
    from the old one; the exact (binomial) version is used for small discordant counts.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_complication_stage_pathologist.xlsx
  >> SCENARIO (narration):
    We measure the agreement of two pathologists assigning the same cases to postoperative complication
    stage. For categorical agreement, kappa is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: pathologist_a
    - 2nd measure / group: pathologist_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Kappa = 0.671 (substantial agreement)

>> COMMENTARY (narration):
    We measured the agreement of two pathologists (pathologist_a vs pathologist_b) in assigning the same cases to the
    same postoperative complication stage -- excluding chance agreement: kappa = 0.67, substantial. Unlike raw percent agreement, kappa
    reports categorical agreement after removing the chance-agreement share, so it is more honest. In medicine it is the
    standard index for measuring how consistent two pathologists'/radiologists' classifications (stage, grade) are.

====================================================================================

#26  Cochran-Mantel-Haenszel
    file: 26_cmh_treatment_stroke_age.xlsx
  >> SCENARIO (narration):
    We test the treatment-stroke relationship controlling for age-group strata. For
    a stratum-controlled relationship, CMH is appropriate.
  >> VARIABLE SELECTION:
    - Variables: treatment
    - Variables: stroke
    - Stratum: age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi^2(1) = 12.62   p < .001 ***   Common OR (MH) = 3.80 [1.75, 8.25]
    Breslow-Day p = 0.787 (OR homogeneous)   H0 REJECTED

>> COMMENTARY (narration):
    We tested the treatment-stroke relationship while controlling for age-group strata: the common Odds Ratio across
    strata = 3.80 [1.75, 8.25], p < .001; Breslow-Day p = 0.79 means this relationship is consistent across all age
    strata. CMH measures the pure strength of the association by holding a confounding stratum variable constant
    (preventing Simpson's paradox). In medicine it is ideal for robustly estimating a treatment-outcome relationship
    while controlling age differences.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_sex_KAH_diabetes.xlsx
  >> SCENARIO (narration):
    We model the joint relationship structure of three categorical variables. For
    more than two categorical dimensions, log-linear is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sex
    - Variables: KAH
    - Variables: diabetes

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 50.54   Pearson chi^2 = 0.456   joint relationship of 3 categorical variables (sex, KAH, diabetes) modeled

>> COMMENTARY (narration):
    We examined the joint relationship structure of three categorical variables (sex, coronary artery disease, diabetes)
    with a log-linear model. The model explains cell frequencies via main effects and interactions; it reveals which
    pairs/triples of variables vary together. It is the multivariable generalization of the two-way cross-tab. In
    medicine it is used to analyze the joint dependency pattern of more than three categorical dimensions (sex x CAD x
    diabetes) -- comorbidity co-occurrence -- balancing parsimony with AIC to select the most explanatory structure.

====================================================================================

#28  Cross-Tabulation
    file: 28_cross_age_symptom.xlsx
  >> SCENARIO (narration):
    We examine age group and symptom in a cross-tab. For the relationship of two
    qualitative variables, cross-tab is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age_group
    - Variables: symptom

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(6) = 27.01   p < .001 ***   Cramer's V = 0.196 (Moderate)   H0 REJECTED

>> COMMENTARY (narration):
    We examined two categorical variables (age group x symptom) in a cross-tab: chi^2(6) = 27.01, p < .001, V = 0.20 --
    a significant, moderate relationship. Cross-tab + chi-square shows the direction and strength of the association
    between two qualitative variables at the cell level. In medicine it is used to describe age-symptom,
    demography-finding relationships; Cramer's V (0.20) measures the practical strength.

====================================================================================

#29  Multiple Response - Frequency
    file: 29_mr_frequency_comorbidity.xlsx
  >> SCENARIO (narration):
    We analyze a multi-select comorbidity question. For a multi-select question,
    multiple-response frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: comorbidity (coklu yanit / multi-response)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)   multi-select "comorbidity" question

>> COMMENTARY (narration):
    We analyzed a question where several options could be checked ("which comorbidities are present?"): each of 250
    patients listed one or more comorbidities. Multiple-response frequency analysis gives the count of checks per option
    and both the response and case percentages separately (percentages sum to over 100, because a patient can have
    several comorbidities). In medicine it is the standard method for correctly summarizing the comorbidity profiles of
    multi-morbid patients (co-existing diseases).

====================================================================================

#30  Multiple Response - Cross-Tab
    file: 30_mr_categorical_comorbidity_age.xlsx
  >> SCENARIO (narration):
    We cross-tabulate the multi-response comorbidity question by age group. To break
    a multi-select by a category, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: comorbidity (coklu yanit / multi-response)
    - Grouping (categorical): age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: age_group   |   multi-response "comorbidity" x age group

>> COMMENTARY (narration):
    We cross-tabulated the multi-response comorbidity question by a single category (age group): we compared the
    per-comorbidity occurrence rates for each age group. A multiple-response cross-tab answers "does one group carry
    certain comorbidities more often than another?". In medicine it is used to compare the comorbidity profiles of
    age/sex groups.

====================================================================================

#31  Multiple Response x Multiple Response
    file: 31_mr_mr_anesthetic_yanetki.xlsx
  >> SCENARIO (narration):
    We cross-tabulate two multi-response questions (anesthetic agents x side effects) against
    each other. For many-to-many co-occurrence, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: anesthetics (coklu / multi)
    - Variables: side_effects (coklu / multi)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180   two multi-responses (anesthetic agents x side_effects) cross-tabulated

>> COMMENTARY (narration):
    We cross-tabulated two separate multi-response questions (anesthetic agents used x side effects) against each other. This is the
    most complex form of tabulation: on both axes a patient contributes to more than one cell. Multiple-response by
    multiple-response reveals many-to-many co-occurrences such as "which anesthetic agents appear together with which side effects?".
    In medicine it is valuable for examining anesthetic-side effect relationships (pharmacovigilance).

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q_symptom_time.xlsx
  >> SCENARIO (narration):
    We test whether a binary symptom status varies across time points in the same
    patients. For 3+ repeated binary measures, Cochran's Q is appropriate.
  >> VARIABLE SELECTION:
    - Columns: zaman-bazli ikili sutunlar / time-wise binary columns

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000 ns   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the binary symptom status (present/absent) measured at different time points (time) in the same
    patients varies significantly from time to time: Q = 0, p = 1.00 -- no difference across time points. Cochran's Q is
    the generalization of McNemar to more than two repeated conditions; it compares 3+ binary measures in the same
    patient. In medicine it is the right method for comparing whether a symptom persists across different times in the
    same patients.

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5klinik.xlsx
  >> SCENARIO (narration):
    We examine all pairwise correlations among five clinical variables. For
    relationship direction and strength, the correlation matrix is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: BMI
    - Variables: MAP_mmHg
    - Variables: glucose_mg_dL
    - Variables: EtCO2_mmHg

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables: age, BMI, MAP_mmHg, glucose_mg_dL, end-tidal CO2_mg_dL
    Strong relationship (|r| >= 0.75): NONE above this threshold

>> COMMENTARY (narration):
    We computed all pairwise Pearson correlations among five clinical variables (age, BMI, mean arterial pressure, glucose,
    end-tidal CO2): none exceeded the 0.75 strong threshold -- the variables move largely independently. This too is a
    valuable finding: most clinical measures carry different information, none substitutes for another. The correlation
    matrix summarizes the direction and strength of relationships at a glance; in medicine it is the first step in
    spotting overlap among indicators (multicollinearity risk) and seeing which variables truly give separate signals.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Agreement
    file: 34_bland_altman_MAP_method.xlsx
  >> SCENARIO (narration):
    We examine the agreement of two methods (manual/automatic) measuring the same
    arterial pressure. For method interchangeability, Bland-Altman is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: bp_manual
    - 2nd measure / group: bp_automatic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Manual (bp_manual): mean = 128.60   |   Automatic (bp_automatic): mean = 127.33
    Bias (mean difference) ~ 1.27 mmHg   |   limits of agreement (LoA) computed

>> COMMENTARY (narration):
    We examined how well two methods measuring the same arterial pressure (manual vs automatic cuff) agree: the mean
    systematic difference (bias) is ~1.3 mmHg, and 95% limits of agreement were reported. Unlike correlation,
    Bland-Altman answers "can the two methods be used interchangeably?" -- high correlation does not mean agreement,
    there may be a systematic shift. In medicine it is the standard method for testing the interchangeability of two
    measuring devices/methods.

====================================================================================

#35  Effect Size
    file: 35_effect_size_antidepresan.xlsx
  >> SCENARIO (narration):
    We measure the practical size of the HAMD-change difference between two
    antidepressant groups. For importance beyond the p-value, effect size is
    appropriate.
  >> VARIABLE SELECTION:
    - Test variable: pain_VAS_change
    - Grouping (categorical): anesthetic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.673 (medium effect)   HAMD-change difference between two anesthetic groups (anesthetic)

>> COMMENTARY (narration):
    We measured the practical size of the HAMD (depression scale) change difference between two antidepressant groups,
    independent of the p-value, with Cohen's d: d = -0.67, a medium effect. The p-value answers "is there a
    difference?"; effect size answers "how important is it?". Because large samples can produce significant but trivial
    differences, reporting effect size is essential. In medicine this measure clarifies whether the difference between
    two anesthetic agents/treatments is clinically noteworthy.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_lab_clinical.xlsx
  >> SCENARIO (narration):
    We examine the joint structure between the laboratory-findings set and the ICU
    clinical-scores set. For the relationship between two multivariate sets, CCA is
    appropriate.
  >> VARIABLE SELECTION:
    - X variables: WBC_K_uL
    - X variables: procalcitonin_ng_mL
    - X variables: creatinine_mg_dL
    - X variables: AST_U_L
    - X variables: Hb_g_dL
    - Y variables: APACHE_score
    - Y variables: SOFA_score
    - Y variables: SAPS_score
    - Y variables: GCS

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.970   chi^2(20) = 866.35   p < .001 ***
    CC2: r = 0.936   chi^2(12) = 457.62   p < .001 ***
    Set-X (lab): WBC, CRP, creatinine, AST, Hb  |  Set-Y (clinical scores): APACHE, SOFA, SAPS, GCS

>> COMMENTARY (narration):
    We resolved the joint structure between two multivariate measure sets -- laboratory findings (WBC, CRP, creatinine,
    AST, Hb) and ICU clinical scores (APACHE, SOFA, SAPS, GCS) -- with canonical correlation. The first two canonical
    functions are very strong (r = 0.970 and 0.936, p < .001): the two sets are intensely related. CCA answers "how is
    one variable set related to another?" in a single step -- it is the multivariate-on-both-sides version of multiple
    regression. In medicine it reveals the latent relationship structure between the laboratory bundle and the clinical
    score bundle.

====================================================================================

#37  Correspondence Analysis
    file: 37_ca_diagnosis_age.xlsx
  >> SCENARIO (narration):
    We map the relationship between diagnosis and age group. To see the structure of
    two categorical variables, CA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: diagnosis
    - Variables: age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.046   (2 dimensions)   diagnosis x age_group

>> COMMENTARY (narration):
    We mapped the relationship in the cross-tab of two categorical variables (diagnosis x age group) into a visual space
    with correspondence analysis: total inertia 0.046 (weak relationship) resolved over two dimensions. CA positions the
    chi-square relationship in a two-dimensional space, showing which categories are close (co-occurring). In medicine
    it is powerful for visually interpreting the structure of qualitative relationships such as diagnosis-age,
    diagnosis-sex; the low inertia indicates the relationship is weak.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_SF36_18.xlsx
  >> SCENARIO (narration):
    We cluster eighteen quality-of-life items (physical/mental/social) by their
    similarity. To find the latent dimension structure, VarClus is appropriate.
  >> VARIABLE SELECTION:
    - Variables: physical_m1..6, mental_m1..6, social_m1..6 (18 madde / items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: physical_m1..6, mental_m1..6, social_m1..6 -- SF-36-like)

>> COMMENTARY (narration):
    We clustered eighteen quality-of-life items (physical, mental, social -- SF-36-style, 6 each) by how related they
    are: they grouped into 3 main dimensions -- most likely matching the natural physical/mental/social domains. VarClus
    groups the VARIABLES, not the observations -- by placing highly correlated items in the same cluster it reveals the
    latent dimensional structure of the data set. In medicine it is practical for reducing a long quality-of-life scale
    to a few core sub-scales and for spotting redundant items.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_bp.xlsx
  >> SCENARIO (narration):
    We model mean arterial pressure with four predictors (age, BMI, salt, exercise). To explain
    a continuous outcome with multiple variables, multiple regression is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: MAP_mmHg
    - Predictor(s): age
    - Predictor(s): BMI
    - Predictor(s): salt_g_day
    - Predictor(s): exercise_hour_week

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.683   Adj. R^2 = 0.676   predictors: age, BMI, salt_g_day, exercise_hour_week

>> COMMENTARY (narration):
    We modeled mean arterial pressure with four predictors (age, BMI, daily salt, exercise) at once: the model explains
    68.3% of variance (Adj. R^2 = 0.68) -- strong explanatory power. Multiple regression gives each predictor's pure
    contribution to MAP with the others held constant; thus it answers "which factor really raises MAP?" while
    controlling confounders. In medicine it is the core method for identifying the risk factors that drive a clinical
    indicator and for prediction.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_KAH.xlsx
  >> SCENARIO (narration):
    We model heart attack (yes/no) with four risk factors. For a binary outcome,
    logistic regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: heart_attack
    - Predictor(s): age
    - Predictor(s): EtCO2_mmHg
    - Predictor(s): smoking
    - Predictor(s): MAP_mmHg

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 = 0.095   binary outcome: heart_attack   predictors: age, end-tidal CO2, smoking, MAP_mmHg

>> COMMENTARY (narration):
    We modeled a binary outcome (heart attack yes/no) with four risk factors: the model has weak-to-moderate explanatory
    power (pseudo R^2 = 0.10) and gives each predictor's effect on the odds. Logistic regression replaces linear
    regression when the outcome is binary; coefficients are converted to Odds Ratios to read "how many times does the
    heart-attack odds change per unit increase in this risk factor?". In medicine it is the core model for predicting
    disease/event risk.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Count (Poisson) Regression
    file: 41_poisson_admission_count.xlsx
  >> SCENARIO (narration):
    We model the annual admission count with age and diabetes presence. For a count
    outcome, Poisson regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: annual_admission_count
    - Predictor(s): age
    - Predictor(s): diabetes_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 590.52   Deviance = 228.88   count outcome: annual_admission_count   predictors: age, diabetes_present

>> COMMENTARY (narration):
    We modeled a count variable (annual intensive care unit admission count) with age and diabetes presence. Poisson regression is
    the right model when the outcome is a count (0,1,2,... items); linear regression is unsuitable because it can
    produce negative/fractional predictions. Coefficients give the effect on the count rate. In medicine it is used to
    explain count outcomes such as admission count, attack count, event frequency with patient features.

====================================================================================

#42  Multinomial Logistic
    file: 42_multinomial_complication_stage.xlsx
  >> SCENARIO (narration):
    We model the multi-category postoperative complication stage with two predictors. For a nominal
    multi-class outcome, multinomial logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: cancer_stage
    - Predictor(s): age
    - Predictor(s): biomarker

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 468.02   Reference class: 'advanced'   outcome: postoperative complication_stage (>2 categories)   predictors: age, sepsis biomarker

>> COMMENTARY (narration):
    We modeled a more-than-two-category outcome (postoperative complication stage) with two predictors (age, sepsis biomarker). Multinomial
    logistic compares each stage against a reference class (here 'advanced') with a separate logistic equation;
    coefficients are read as "as X increases, how does the chance of being in this stage change relative to the
    reference?". When the outcome is nominal/ordinal with more than two classes it is the right choice. In medicine it
    is used for multi-stage/multi-category diagnostic classification.

====================================================================================

#43  Ordinal Logistic
    file: 43_ordinal_response_level.xlsx
  >> SCENARIO (narration):
    We model the ordinal response level with dose and age. For an ordinal outcome,
    ordinal logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: response_level
    - Predictor(s): propofol_dose_mg
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 344.47   ordinal outcome: response_level   predictors: dose_mg, age

>> COMMENTARY (narration):
    We modeled an ordinal outcome (treatment response level: low/medium/high) with dose and age. Ordinal logistic uses
    the ORDER information between categories (which multinomial ignores); with a "proportional odds" assumption it
    explains all thresholds with one coefficient set. In medicine it is the right and more powerful choice for modeling
    naturally ordered outcomes such as response level, stage, grade.

====================================================================================

#44  PLS Regression
    file: 44_pls_gene_progression.xlsx
  >> SCENARIO (narration):
    We predict a disease progression score from 12 gene expressions. For many highly
    correlated predictors, PLS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: progression_score
    - Predictor(s): gene_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 (train) = 0.894   R^2 (5-fold CV) = 0.880   outcome: progression_score   predictors: gene_01..12 (12 genes)

>> COMMENTARY (narration):
    We predicted a disease progression score from 12 gene expressions. Because the gene expressions are highly
    correlated (multicollinearity), classic regression becomes unstable; PLS reduces them to a few latent components and
    regresses on those. With cross-validated R^2 = 0.88, the model is both strong and generalizable. PLS is ideal when
    predictors are numerous or highly correlated; in genomic/omics data (predicting an outcome from many genes) it is
    widely used.

====================================================================================

#45  Probit Regression
    file: 45_probit_anesthetic_dose_response.xlsx
  >> SCENARIO (narration):
    We estimate the probability of responding to an anesthetic from dose with a probit
    model. For a binary dose-response outcome, probit is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: response
    - Predictor(s): propofol_dose_mg

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 118.36   Pseudo R^2 (McFadden) = 0.535   binary outcome: response   predictor: dose_mg

>> COMMENTARY (narration):
    We estimated the dose-response relationship -- the probability of responding to an anesthetic from dose (dose_mg) -- with a
    probit model: the model is very strong (pseudo R^2 = 0.53). Probit, like logistic, applies to binary outcomes; the
    difference is that its link function is the normal distribution. Probit is classic in dose-response (pharmacology)
    studies (ED50 estimation). In medicine it is a robust choice for modeling a binary clinical response against dose.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_hastane_expenditure.xlsx
  >> SCENARIO (narration):
    We model a threshold-piled out-of-pocket health-expenditure variable with
    admission days and insurance presence. For a censored outcome, tobit is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: pocket_expenditure_TL
    - Predictor(s): admission_day
    - Predictor(s): insurance_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 2868.75   Left censor   outcome: pocket_expenditure_TL (censored)   predictors: admission_day, insurance_present

>> COMMENTARY (narration):
    We modeled a lower-bounded (threshold-piled) out-of-pocket health-expenditure variable with admission days and
    insurance presence. Tobit is for "censored" dependent variables that pile up at a threshold; ordinary regression
    gives biased estimates by ignoring this pile-up. In medicine/health economics it is the right model for floor-effect
    outcomes (zero out-of-pocket spending); coefficients reflect the true (uncensored) relationship.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_dose_response.xlsx
  >> SCENARIO (narration):
    We model the treatment response score with dose and age in a Bayesian framework.
    To express uncertainty probabilistically, Bayesian regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: response_score
    - Predictor(s): propofol_dose_mg
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2 posterior mean = 27.25 (sd = 3.89)   outcome: response_score   predictors: dose_mg, age
    coefficient posteriors + P(beta>0) reported

>> COMMENTARY (narration):
    We modeled the treatment response score with dose and age in a Bayesian framework: instead of point estimates we
    obtained each coefficient's full posterior distribution and the "probability the effect is positive". The Bayesian
    approach expresses uncertainty directly in probability language and can incorporate prior knowledge. In anesthesia and critical care,
    when the sample is small or prior-study information is valuable, it offers intuitive interpretations like "the dose
    effect is probably positive".

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear_KMY_age.xlsx
  >> SCENARIO (narration):
    We model the S-shaped relationship of bone mineral density with age via a
    logistic curve. For a curved relationship, nonlinear regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: BIS_index
    - Predictor(s): age_year
    - Value: fonksiyon/function: logistic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Function: y = K / (1 + exp(-r*(x-x0)))  (logistic growth)   R^2 = 0.960

>> COMMENTARY (narration):
    We modeled the relationship of bone mineral density (BMD) with age not as a straight line but as an S-shaped
    logistic curve: the fit is very high (R^2 = 0.96). Nonlinear regression fits a theoretical function form directly to
    the data when the relationship is curved (development, saturation, threshold) and makes the parameters (ceiling K,
    rate r, inflection x0) interpretable. In medicine it is the right tool for modeling age-related physiological change
    curves (bone density, growth).

====================================================================================

#49  Ridge Regression
    file: 49_ridge_biomarker_progression.xlsx
  >> SCENARIO (narration):
    We predict disease progression from 15 correlated sepsis biomarkers with ridge. For
    multicollinearity, ridge is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: progression
    - Predictor(s): biomarker_p01..15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 1.0   R^2 = 0.798   outcome: progression   predictors: sepsis biomarker_p01..15 (15 sepsis biomarkers)

>> COMMENTARY (narration):
    We predicted disease progression from 15 correlated sepsis biomarkers with ridge regression (R^2 = 0.80). Ridge adds an L2
    penalty to shrink all coefficients in magnitude without zeroing them; this prevents the instability caused by high
    correlation among sepsis biomarkers (multicollinearity). Because clinical sepsis biomarkers are inherently correlated, ridge gives
    stable results on such data. In medicine it is preferred for prediction with many co-varying sepsis biomarkers.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso_30_pred_5_active.xlsx
  >> SCENARIO (narration):
    We model treatment response from 30 candidate features with lasso, selecting the
    important ones. For automatic variable selection, lasso is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: treatment_response
    - Predictor(s): x01..30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.1   R^2 = 0.723   18 of 30 features shrunk to zero (automatic variable selection)

>> COMMENTARY (narration):
    We modeled treatment response from 30 candidate features with lasso regression: R^2 = 0.72, and lasso shrank 18 of
    the 30 feature coefficients exactly to zero, selecting only the effective variables. This is its difference from
    ridge: because lasso can zero coefficients, it performs prediction and variable selection at the same time. In
    medicine it is very useful for automatically winnowing the "few truly important factors" out of many candidate
    sepsis biomarkers/features (a sparse, interpretable model).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_stress_cortisol_bp.xlsx
  >> SCENARIO (narration):
    We test the stress -> cortisol -> mean arterial pressure chain. To resolve the intermediate
    mechanism, mediation analysis is appropriate.
  >> VARIABLE SELECTION:
    - Predictor(s): stress_PSS
    - Value: M (araci/mediator): cortisol_ng_mL
    - Dependent variable: MAP_mmHg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = -0.142   95% CI [-0.688, 0.355]   (includes zero; not significant)
    X: stress_PSS  M: cortisol_ng_mL  Y: MAP_mmHg

>> COMMENTARY (narration):
    We tested the chain "stress (X) -> cortisol (M) -> mean arterial pressure (Y)": the indirect effect is -0.142, its 95% CI
    includes zero -- so stress's indirect effect on arterial pressure through cortisol was not significant. Mediation
    analysis resolves "why/how does X affect Y?" through an intermediate mechanism; here this mechanistic path was not
    supported. In medicine it is powerful for understanding which biological intermediate process (cortisol) a factor's
    (stress) effect flows through -- or does not.

====================================================================================

#52  Path Analysis
    file: 52_path_health_behavior.xlsx
  >> SCENARIO (narration):
    We test direct/indirect relationships as a single causal diagram. For a
    relationship network, path analysis is appropriate.
  >> VARIABLE SELECTION:
    - Value: behavior ~ info + attitude + intention
    - Value: intention ~ attitude

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.928   RMSEA = 0.200   model: behavior ~ info + attitude + intention; intention ~ attitude

>> COMMENTARY (narration):
    We tested the direct and indirect relationships among several variables as a single causal diagram. The fit indices
    (CFI = 0.93, RMSEA = 0.20) show the model fits the data reasonably-to-partially. Path analysis estimates the whole
    relationship network at once instead of separate regressions; it shows how variables affect each other and a common
    outcome. In medicine it is used to test health-behavior theories (information -> attitude -> intention -> behavior).

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_patient_visit_SpO2.xlsx
  >> SCENARIO (narration):
    We model SpO2 measured repeatedly across visits in the same patients, taking
    patient as a random effect. For repeated/nested data, LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: SpO2_pct
    - Predictor(s): visit
    - Cluster: patient_id (random)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC = 0.936   Group variance (random intercept) = 1.724   outcome: SpO2   fixed: visit   group: patient_id

>> COMMENTARY (narration):
    We modeled SpO2 measured repeatedly across visits in the same patients, taking patient identity as a random effect.
    ICC = 0.94 is very high: almost all SpO2 variability comes from between-patient differences, while within-patient
    visits are very similar. LMM correctly handles dependency in nested/repeated (measurements within patient) data; it
    solves the "independence" assumption that ordinary regression violates via random effects. In medicine it is the
    right choice for panel/repeated-measure data (patient follow-up).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation_missing_patient.xlsx
  >> SCENARIO (narration):
    Instead of deleting missing data we fill it with 5 plausible value sets. To
    handle missingness without bias, multiple imputation is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: MAP_mmHg
    - Predictor(s): age
    - Predictor(s): BMI
    - Predictor(s): SpO2_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (imputation) = 5   outcome: MAP_mmHg   predictors: age, BMI, SpO2

>> COMMENTARY (narration):
    Instead of deleting missing data, we analyzed by producing 5 plausible value sets (m = 5) and combining the results.
    Multiple imputation -- unlike filling gaps with a single estimate (which ignores uncertainty) -- accounts for
    imputation uncertainty too, yielding unbiased estimates and correct standard errors. In medicine it is the modern
    standard for handling the inevitable gaps in clinical/laboratory data (missing measurement/record) without shrinking
    the sample or distorting results.

====================================================================================

#55  GEE
    file: 55_gee_SpO2_treatment.xlsx
  >> SCENARIO (narration):
    We model SpO2 in repeated visit measures of the same patients. For a
    population-average effect, GEE is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: SpO2_pct
    - Predictor(s): visit
    - Cluster: patient_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coef = -0.098   p < .001 ***   QIC = 306.75   outcome: SpO2   group: patient_id

>> COMMENTARY (narration):
    We modeled SpO2 in repeated visit measures of the same patients; the visit effect is significant and negative
    (b = -0.098, p < .001) -- SpO2 declines over follow-up. GEE estimates the POPULATION-AVERAGE effect rather than
    individual effects in repeated/clustered data and corrects within-group correlation with a "working correlation
    structure". While LMM focuses on individual random effects, GEE focuses on the average trend. In medicine it is
    preferred for population-level questions like "what is the time/treatment effect in the average patient?".

====================================================================================

#56  GLMM
    file: 56_glmm_side_effect.xlsx
  >> SCENARIO (narration):
    We model a repeatedly measured side-effect count in the same patients with month
    and new anesthetic, taking patient as a random effect. For repeated counts, GLMM is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: side_effect_count
    - Predictor(s): time_month
    - Predictor(s): new_anesthetic
    - Cluster: patient_id (Poisson)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    new_anesthetic coef = 0.541   p < .001 ***   outcome: side_effect_count (count, Poisson)   group: patient_id

>> COMMENTARY (narration):
    We modeled a COUNT outcome (side-effect count) measured repeatedly in the same patients with month and new anesthetic,
    taking patient as a random effect; the new-anesthetic effect is significant and positive (b = 0.54, p < .001) -- the new
    anesthetic increases the side-effect count. GLMM extends LMM to non-normal outcomes (count, binary): it handles both the
    distribution (Poisson) and the clustering (random effect) at the same time. In medicine it is the right model for
    repeatedly measured binary/count outcomes (monthly side-effect/attack count).

====================================================================================

#57  Elastic Net
    file: 57_elasticnet_risk_score.xlsx
  >> SCENARIO (narration):
    We model a risk score from 40 predictors with Elastic Net. For many clustered
    predictors, Elastic Net is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: risk_score
    - Predictor(s): x01..40

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.123   L1 ratio = 0.50   R^2 (train) = 0.491   40 predictors

>> COMMENTARY (narration):
    We modeled a risk score from 40 predictors with Elastic Net. Elastic Net blends the ridge (L2) and lasso (L1)
    penalties (L1 ratio = 0.5): it both keeps groups of correlated variables together (ridge property) and zeroes out
    redundant ones (lasso property). It thus provides a balanced model with high-dimensional, clustered predictors. In
    medicine it is chosen when there are many clustered sepsis biomarkers/features, where lasso or ridge alone is insufficient.

====================================================================================

#58  Robust Regression
    file: 58_robust_age_creatinine.xlsx
  >> SCENARIO (narration):
    We model creatinine with age, down-weighting outliers. For data with outliers,
    robust regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: creatinine_mg_dL
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust Intercept = 0.695   OLS Intercept = 0.993   outcome: creatinine_mg_dL   predictor: age

>> COMMENTARY (narration):
    When modeling creatinine with age, we used robust regression to prevent outliers from distorting the estimate. The
    gap between the robust and OLS intercepts (0.695 vs 0.993) shows a few outlying observations (e.g. renal-failure
    patients) pull the classic estimate; the robust method down-weights them to reflect the "typical" relationship. In
    medicine it is the right way to get robust estimates without deleting outliers in laboratory data that contains them.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_hastane_cost.xlsx
  >> SCENARIO (narration):
    We model different points of total treatment cost's distribution (lower 10%,
    median, upper 90%) separately. For varying effects, quantile regression is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: total_cost_TL
    - Predictor(s): age
    - Predictor(s): admission_day
    - Quantiles: 0.10 / 0.50 / 0.90

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    separate slopes reported for q=0.10, q=0.50, q=0.90   outcome: total_cost_TL   predictors: age, admission_day

>> COMMENTARY (narration):
    We modeled not just the mean of total treatment cost but different points of the distribution (lower 10%, median,
    upper 90%) separately. Predictor effects can vary by quantile -- a factor may be weak for low-cost patients and very
    strong for high-cost (expensive) cases. Quantile regression gives the true picture when the "mean effect" is
    misleading (effect varies across the distribution). In medicine/health economics it shows what classic regression
    misses by examining cost inequality and expensive-case behavior.

====================================================================================

#60  ROC Curve
    file: 60_roc_biomarker_complication.xlsx
  >> SCENARIO (narration):
    We assess how well a sepsis biomarker score separates a binary outcome with ROC. For
    discrimination and threshold selection, ROC is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: cancer_present
    - Predictor(s): biomarker_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.969 (excellent discrimination)   Youden optimum threshold = 0.469   Sensitivity = 0.925, 1-Specificity = 0.064
    n = 250 (positive 93, negative 157)

>> COMMENTARY (narration):
    We assessed how well a sepsis biomarker score separates a binary outcome (postoperative complication_present) with a ROC curve: AUC = 0.97,
    excellent discrimination; the optimum decision threshold was set at 0.47 via Youden. ROC shows the
    sensitivity-specificity trade-off at all possible thresholds and evaluates the model without being tied to a single
    threshold. In medicine it is the standard tool for measuring a diagnostic/screening test's discriminative power and
    selecting the best decision threshold (cutoff).

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss_disease_prediction.xlsx
  >> SCENARIO (narration):
    We measure a diagnostic/risk model's discrimination with TSS. For a fair measure
    under class imbalance, TSS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: disease_present
    - Predictor(s): risk_factor1
    - Predictor(s): risk_factor2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.39 (acceptable)   Sensitivity = 0.44   Specificity = 0.95   N = 200
    binary outcome: disease_present   predictors: risk_factor1, risk_factor2

>> COMMENTARY (narration):
    We measured a diagnostic/risk model's (disease present/absent) discrimination with TSS: TSS = 0.39 (sensitivity
    0.44, specificity 0.95). TSS = sensitivity + specificity - 1; it excludes chance-expected success and gives a fair
    performance measure even with imbalanced classes. Here the model is specific (few false positives) but not sensitive
    (tends to miss true cases). While accuracy can be biased, TSS evaluates the power to capture both positives and
    negatives together. In medicine it is robust for reporting the true discriminative power of diagnostic/risk
    classification models.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity_diabetes_class.xlsx
  >> SCENARIO (narration):
    We evaluate a diagnostic classifier's predictions against true diagnoses. For
    detailed performance, the confusion matrix is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: actual_diagnosis
    - 2nd measure / group: prediction_diagnosis

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.94   F1 = 0.923   (true diagnosis vs predicted diagnosis)

>> COMMENTARY (narration):
    We evaluated a diagnostic classifier's predictions against true diagnoses via a confusion matrix: accuracy 94%,
    F1 = 0.92. The confusion matrix gathers true/false positives and negatives in one table, from which sensitivity,
    specificity, precision and F1 are derived. Because a single accuracy number can mislead (especially with imbalanced
    classes in rare diseases), this metric set shows where the model errs. In medicine it is fundamental for detailed
    reporting of diagnostic/prediction model performance.

====================================================================================

#63  Random Forest
    file: 63_rf_complication_type.xlsx
  >> SCENARIO (narration):
    We classify postoperative complication type from five features with a random forest. For nonlinear
    diagnosis, random forest is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: cancer_type
    - Predictor(s): age
    - Predictor(s): sex_E
    - Predictor(s): CA125
    - Predictor(s): creatinine
    - Predictor(s): Hb_g_dL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.900   outcome: postoperative complication_type   predictors: age, sex_E, CA125, creatinine, Hb_g_dL

>> COMMENTARY (narration):
    We classified postoperative complication type from five features with a random forest: accuracy 90%. Random forest combines the votes of
    hundreds of decision trees; it automatically captures nonlinear relationships and interactions and also gives a
    variable-importance ranking. It reduces a single tree's overfitting by averaging. In medicine it is widely used for
    complex, nonlinear diagnostic/prediction problems (postoperative complication type, phenotype) because it offers both high accuracy and
    "which sepsis biomarker matters" information.

====================================================================================

#64  Support Vector Machines (SVM)
    file: 64_svm_sepsis.xlsx
  >> SCENARIO (narration):
    We classify sepsis present/absent from four clinical features with SVM. For
    well-separable classes, SVM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: sepsis
    - Predictor(s): procalcitonin_ng_mL
    - Predictor(s): WBC_K_uL
    - Predictor(s): core_temp_C
    - Predictor(s): lactate_mmol

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: sepsis   predictors: CRP_mg_L, WBC_K_uL, fever_C, lactate_mmol

>> COMMENTARY (narration):
    We classified sepsis present/absent from four clinical features with SVM: accuracy 100% -- the classes are perfectly
    separable with these markers (very high accuracy should be confirmed against overfitting via cross-validation). SVM
    finds the decision boundary separating classes with the widest margin; with the kernel trick it can also do
    nonlinear separation. In medicine it is a strong, stable classifier for well-separable diagnostic problems (sepsis,
    acute conditions), especially on small-to-medium data.

====================================================================================

#65  Gradient Boosting
    file: 65_gb_repeat_admission.xlsx
  >> SCENARIO (narration):
    We classify repeat-admission risk from four predictors with gradient boosting.
    For top accuracy, gradient boosting is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: repeat_admission_risk
    - Predictor(s): age
    - Predictor(s): comorbidity_count
    - Predictor(s): last_admission_day
    - Predictor(s): anesthetic_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.893   outcome: repeat_admission_risk   predictors: age, comorbidity_count, last_admission_day, anesthetic_count

>> COMMENTARY (narration):
    We classified repeat-admission risk (repeat_admission_risk) from four predictors with gradient boosting: accuracy
    89.3%. Gradient boosting adds weak trees sequentially -- each new tree corrects the previous model's errors; this is
    why it wins most prediction competitions. While random forest votes in parallel, boosting reduces error step by
    step. In medicine it is preferred where the highest predictive accuracy is sought (readmission, complication-risk
    scoring); it requires tuning against overfitting.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_patient_4kume.xlsx
  >> SCENARIO (narration):
    We cluster patients by four features. For patient phenotyping, k-means is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: BMI
    - Variables: MAP_mmHg
    - Variables: glucose
    - Number of clusters: 4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.442   variables: age, BMI, MAP_mmHg, glucose

>> COMMENTARY (narration):
    We split patients into 4 clusters by four features (age, BMI, mean arterial pressure, glucose): silhouette = 0.44, reasonable-
    moderate separation. K-means assigns observations to the nearest cluster center and iteratively updates the centers;
    it groups similar patients into natural clusters. No labels are needed (unsupervised). In medicine it is the core
    method for patient phenotyping (risk groups, metabolic profiles) and patient segmentation; silhouette checks the
    appropriateness of the cluster count.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5tip_10gen.xlsx
  >> SCENARIO (narration):
    We hierarchically cluster patients by ten gene expressions. For nested subtype
    structure, hierarchical clustering is appropriate.
  >> VARIABLE SELECTION:
    - Variables: gene_01..10
    - Number of clusters: 5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 5   Silhouette = 0.147   variables: gene_01..10 (10 genes)

>> COMMENTARY (narration):
    We split patients into 5 groups by ten gene expressions with hierarchical clustering (silhouette = 0.15, weak
    separation -- groups partly overlap). Hierarchical clustering merges observations step by step to build a tree
    (dendrogram); its difference from k-means is not having to fix the cluster count in advance and seeing the nested
    structure. In medicine/genomics it is used to explore a hierarchy of patient or gene groups (subtypes) and to read
    the natural cluster count from the dendrogram.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan_case_location.xlsx
  >> SCENARIO (narration):
    We cluster case locations with density-based DBSCAN. For spatial clusters and
    outliers, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Variables: lat
    - Variables: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.911   variables: lat, lon (location)

>> COMMENTARY (narration):
    We clustered case locations (lat, lon) with density-based DBSCAN: 4 dense clusters, silhouette = 0.91, very good
    separation. Unlike k-means, DBSCAN does not require the cluster count in advance, can find clusters of any shape,
    and marks sparse points as "noise". In medicine/epidemiology it is ideal for detecting spatial case concentrations
    (outbreak clusters, disease foci) and isolating outlier locations.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca_clinical_6ozellik.xlsx
  >> SCENARIO (narration):
    We reduce six correlated clinical indicators to a few components with PCA. For
    dimensionality reduction, PCA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: BMI
    - Variables: MAP_mmHg
    - Variables: glucose
    - Variables: EtCO2_mmHg
    - Variables: SpO2_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 explains 91.9% of variance   variables: age, BMI, MAP_mmHg, glucose, end-tidal CO2, SpO2

>> COMMENTARY (narration):
    We reduced six correlated clinical indicators to a few components with PCA: the first component alone explains 91.9%
    of variance -- so these six measures largely reflect a single latent dimension (general metabolic/cardiovascular
    risk). PCA transforms correlated variables into mutually independent components; it reduces dimensions, eases
    visualization and resolves multicollinearity. In medicine it is fundamental for summarizing multi-indicator clinical
    data and building a "composite risk index".

====================================================================================

#70  t-SNE
    file: 70_tsne_complication_5tip.xlsx
  >> SCENARIO (narration):
    We reduce 12-gene-dimensional data to 2 dimensions with t-SNE for visualization.
    For nonlinear subtype discovery, t-SNE is appropriate.
  >> VARIABLE SELECTION:
    - Variables: gene_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence = 0.558 (good)   high-dimensional 12 genes embedded into 2 dimensions

>> COMMENTARY (narration):
    We reduced 12-gene-dimensional data to 2 dimensions for visualization with t-SNE (KL = 0.56, good quality). t-SNE
    tries to preserve high-dimensional neighborhoods, placing similar observations near and dissimilar ones far; it
    reveals nonlinear cluster structures visually that PCA misses. Interpretation is visual (the axes have no absolute
    meaning). In medicine/genomics it is used for exploratory visualization of hidden patient subtypes in
    high-dimensional gene/omics data.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds_patient_distance.xlsx
  >> SCENARIO (narration):
    We place the inter-patient similarity structure onto a 2D map. For similarity
    maps, MDS is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: BMI
    - Variables: MAP_mmHg
    - Variables: EtCO2_mmHg
    - Variables: SpO2_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.091 (acceptable)   variables: age, BMI, MAP_mmHg, end-tidal CO2, SpO2

>> COMMENTARY (narration):
    We placed the distance/similarity structure among patients onto a 2-dimensional map with acceptable stress (0.09):
    low stress means the map represents the true distances well. MDS positions observations in an interpretable space by
    preserving inter-observation distances; its difference from t-SNE is the aim of preserving global distance structure.
    In medicine it is used to draw patient/group similarity maps and for clinical-profile positioning.

====================================================================================

#72  UMAP
    file: 72_umap_omics_5tip.xlsx
  >> SCENARIO (narration):
    We reduce 25-dimensional omics data to 2 dimensions with UMAP. For both local
    and global structure, UMAP is appropriate.
  >> VARIABLE SELECTION:
    - Variables: omics_01..25

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    25 omics variables embedded into 2 dimensions   (local-global balance via n_neighbors / min_dist)

>> COMMENTARY (narration):
    We reduced 25-dimensional omics (gene/protein) data to 2 dimensions with UMAP. UMAP does nonlinear dimensionality
    reduction like t-SNE but preserves both local and global structure better and is faster. Larger n_neighbors
    emphasizes broader groups, larger min_dist emphasizes the gaps between clusters. In medicine/genomics it is a modern
    choice for visualizing high-dimensional omics data and exploring natural patient subtypes.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_BDI_21.xlsx
  >> SCENARIO (narration):
    We measure the internal consistency of a 21-item scale. For scale reliability,
    Cronbach's alpha is appropriate.
  >> VARIABLE SELECTION:
    - Variables: item_01..21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach alpha = 0.974 (excellent internal consistency)   21 items

>> COMMENTARY (narration):
    We measured the internal consistency of a 21-item scale (e.g. a quality-of-life / patient-reported outcome) with
    Cronbach's alpha: alpha = 0.97, excellent -- the items measure the same construct consistently. Cronbach's alpha
    shows how much the items of a scale "move together"; above 0.70 is considered acceptable. (A very high value can
    also signal item redundancy.) In medicine it is the standard index for reporting the reliability of patient-reported
    scales.

====================================================================================

#74  Likert Scale Analysis
    file: 74_likert_survival_quality_3boyut.xlsx
  >> SCENARIO (narration):
    We analyze a 15-item Likert set of three sub-scales (physical/mental/social).
    For ordinal scale summary and reliability, Likert analysis is appropriate.
  >> VARIABLE SELECTION:
    - Variables: physical_1..5 / mental_1..5 / social_1..5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of items k = 15   Cronbach alpha = 0.782 (acceptable)

>> COMMENTARY (narration):
    We analyzed a 15-item Likert set of three sub-scales (physical, mental, social): reliability is acceptable
    (alpha = 0.78), and item distributions and central tendencies were reported. Likert analysis describes ordinal scale
    responses (1-5) with appropriate summaries (median, distribution, pile-up) and checks scale reliability. In medicine
    it is used for correctly summarizing and interpreting quality-of-life/patient-reported outcome scales.

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18_item_3faktor.xlsx
  >> SCENARIO (narration):
    We discover the latent factors behind eighteen items. For scale-structure
    discovery, EFA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: q01..18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO = 0.909 (excellent)   factor analysis appropriate   18 items (q01..18)

>> COMMENTARY (narration):
    We applied EFA to discover how many latent factors lie behind eighteen items: KMO = 0.91, the data is very suitable
    for factor analysis. EFA reduces observed items to a few unobserved "factors"; it reveals which items measure the
    same dimension. In medicine it is fundamental when developing a new patient-reported scale (discovering structure)
    and mapping item groups to theoretical domains; KMO and Bartlett are prerequisite tests.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3radyolog_tumor.xlsx
  >> SCENARIO (narration):
    We compare the value three radiologists measured on the same images with ICC.
    For inter-rater reliability on continuous measures, ICC is appropriate.
  >> VARIABLE SELECTION:
    - Variables: radiologist_1_mm
    - Variables: radiologist_2_mm
    - Variables: radiologist_3_mm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.951 (Excellent)   95% CI [0.920, 0.970]   3 radiologists (radiologist_1/2/3_mm)

>> COMMENTARY (narration):
    We compared the value (mm) three radiologists measured on the same images with ICC: ICC = 0.95, excellent -- the
    radiologists measure almost identically. Unlike kappa, ICC measures inter-rater reliability on CONTINUOUS measures
    and can assess both consistency and absolute agreement. In medicine it is the standard index for determining how
    reliable/interchangeable different experts' measurements (radiology/pathology) are.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_DASS21_12.xlsx
  >> SCENARIO (narration):
    We test whether a predefined three-factor structure (depression/anxiety/stress
    -- DASS) fits the data. For construct validity, CFA is appropriate.
  >> VARIABLE SELECTION:
    - Value: depression: depression_1..4
    - Value: anxiety: anxiety_1..4
    - Value: stress: stress_1..4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 1.002   RMSEA = 0.000   3 factors: depression / anxiety / stress (4 items each -- DASS-like)

>> COMMENTARY (narration):
    Unlike EFA, we TESTED whether a pre-defined three-factor structure (depression/anxiety/stress -- DASS-style) fits
    the data with CFA: the fit indices are excellent (CFI = 1.00, RMSEA = 0.00). CFA tests a theoretical scale model --
    which item loads on which factor is fixed in advance, and the question is "does the model fit the data?". In medicine
    it is a mandatory step for confirming the construct validity of a developed mental-health scale (do the DASS
    dimensions match theory).

====================================================================================

#78  Survey Mean
    file: 78_survey_means_bp.xlsx
  >> SCENARIO (narration):
    We estimate the mean mean arterial pressure in a stratified/weighted health survey. For a
    complex sample, design-based mean is appropriate.
  >> VARIABLE SELECTION:
    - Variables: MAP_mmHg
    - Weight: weight
    - Stratum: age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    MAP_mmHg: M_hat = 129.09   SE = 0.79   95% CI [127.53, 130.65]   CV 0.62%   weight: weight, stratum: age_group

>> COMMENTARY (narration):
    In a stratified/weighted health-survey design we estimated the mean mean arterial pressure while accounting for the design:
    M = 129.09 mmHg, with a Taylor-linearization SE and 95% CI [127.5, 130.7]. Complex-sample methods account for
    unequal selection probabilities (weights) and stratification; ignoring these biases the standard errors. In medicine
    it is the right way to produce correct point estimates and confidence intervals from national/regional representative
    health surveys (NHANES-style).

====================================================================================

#79  Survey Frequency
    file: 79_survey_freq_smoking.xlsx
  >> SCENARIO (narration):
    We estimate smoking-status category proportions accounting for the design. For a
    complex sample, design-based frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: smoking
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    weighted proportions + Taylor SE + 95% CI for smoking categories
    (e.g. 'active low' p_hat = 0.185, 'active high' p_hat = 0.042)

>> COMMENTARY (narration):
    We estimated the category proportions of a categorical survey question (smoking status) under sampling weights and
    stratification; each proportion has a design-based SE and confidence interval. Complex-survey frequency analysis,
    unlike a simple percentage, estimates population proportions without bias by accounting for the sampling design. In
    medicine it is used to correctly report risk-factor (smoking) distributions from representative health surveys.

====================================================================================

#80  Survey Total
    file: 80_survey_total_diabetes_prediction.xlsx
  >> SCENARIO (narration):
    We estimate the population total (total diabetes patients) from the sample. For
    a complex sample, design-based total is appropriate.
  >> VARIABLE SELECTION:
    - Variables: diabetes_patient_count
    - Weight: weight
    - Stratum: size_stratum

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    diabetes_patient_count: T_hat = 727,848   SE = 7,969.09   95% CI [712,160, 743,536]   stratum: size_stratum

>> COMMENTARY (narration):
    We estimated the population TOTAL (total diabetes patient count) from the sample: T = 727,848, 95% CI [712,160,
    743,536]. Weights tell how many population units each observation represents; the total is estimated by summing
    those weights, with uncertainty reported via Taylor SE. In medicine/public health it is the right method for
    producing population-scaled totals (total patients, total burden) from a sample.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg_bp.xlsx
  >> SCENARIO (narration):
    We regress mean arterial pressure accounting for the design. For relationships in complex
    samples, design-based regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: MAP_mmHg
    - Predictor(s): age
    - Predictor(s): BMI
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.547   age coef = 0.473 (p < .001)   outcome: MAP_mmHg   predictors: age, BMI   weight: weight

>> COMMENTARY (narration):
    We regressed mean arterial pressure while accounting for the sampling design (weight, stratum): age is a
    significant predictor (b = 0.47, p < .001), and the model explains 54.7% of variance. Design-based regression
    incorporates weights and the cluster/stratum structure into coefficient and standard-error calculation; ordinary
    regression ignores these and gives biased inference. In medicine it is the right way to model relationships in
    representative health-survey data.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic_diabetes.xlsx
  >> SCENARIO (narration):
    We model diabetes with design-weighted logistic. For binary outcomes in complex
    samples, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: diabetes
    - Predictor(s): age
    - Predictor(s): BMI
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    age: OR = 1.051   p < .001 ***   binary outcome: diabetes   predictors: age, BMI   weight: weight

>> COMMENTARY (narration):
    We modeled a binary outcome (diabetes yes/no) with design-weighted logistic regression: each one-unit increase in age
    raises the diabetes odds by ~5.1% (OR = 1.051, p < .001) -- a strong predictor. This method extends logistic
    regression to complex sample designs -- weight and stratum are reflected in the standard errors. In medicine it is
    used to correctly estimate the probability of a binary outcome (disease present/absent) from representative health
    surveys.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_mortality_age.xlsx
  >> SCENARIO (narration):
    We model 1-year mortality probability with age via a flexible curve. For a
    nonlinear effect, GAM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: 1yr_mortality_probability
    - Predictor(s): age (smooth)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 (explained) = 0.983   outcome: 1yr_mortality_probability   predictor: age (smooth term)

>> COMMENTARY (narration):
    We modeled 1-year mortality probability with age, without assuming a straight line in advance, via a flexible curve
    (smooth): the explained variance is very high (pseudo R^2 = 0.98). GAM extends linear regression -- it models each
    predictor's effect as a smooth function learned from the data, capturing curved relationships without losing
    interpretability. In medicine it is a more explanatory choice than black-box models for relationships where the
    effect is nonlinear (mortality rising exponentially with age).

====================================================================================

#84  Discriminant Analysis
    file: 84_diskriminant_3sinif_5lab.xlsx
  >> SCENARIO (narration):
    We classify patient class from five laboratory measures. To assign to predefined
    diagnostic classes, discriminant analysis is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: class
    - Predictor(s): WBC_K_uL
    - Predictor(s): procalcitonin_ng_mL
    - Predictor(s): creatinine
    - Predictor(s): AST_U_L
    - Predictor(s): Hb_g_dL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: class   predictors: WBC_K_uL, CRP_mg_L, creatinine, AST_U_L, Hb_g_dL

>> COMMENTARY (narration):
    We classified patient class from five laboratory measures with discriminant analysis: accuracy 100% -- the classes
    are fully separable with these markers (overfitting risk should be checked via cross-validation). Discriminant
    analysis finds the linear combinations that best separate groups; it both classifies and shows which marker is most
    influential in separation. In medicine it is used to assign new patients to predefined diagnostic classes and to
    identify discriminating laboratory markers.

====================================================================================

#85  Conditional Logit
    file: 85_conditional_logit_treatment.xlsx
  >> SCENARIO (narration):
    We model patients' choices among treatment alternatives by option features. For
    discrete-choice data, conditional logit is appropriate.
  >> VARIABLE SELECTION:
    - Chooser: patient_id
    - Choice: chosen
    - Alternative features: duration_month
    - Alternative features: cost_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McFadden pseudo R^2 = 1.00   chooser: patient_id   choice: chosen   features: duration_month, cost_TL

>> COMMENTARY (narration):
    We modeled the choices patients made among treatment alternatives by the alternatives' features (duration, cost)
    with a conditional logit. This model is for "discrete choice" data where each individual picks one from a choice set;
    it estimates how an option's features affect its probability of being chosen. Its difference from standard logistic
    is that the choice is conditional on the individual's option set. In medicine it is the core method for treatment/
    service preference modeling (duration-cost effect, patient preference).

====================================================================================

#86  Kaplan-Meier Survival
    file: 86_km_complication_survival.xlsx
  >> SCENARIO (narration):
    We examine patients' time to death and stage differences. For censored time
    data, Kaplan-Meier is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death
    - Grouping (categorical): stage

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 41.1 months   Log-rank chi^2(2) = 1.00, p = 0.607 ns   time: duration_month, event: death, group: stage

>> COMMENTARY (narration):
    We examined the time patients take to reach death with Kaplan-Meier: median survival 41.1 months; stage survival
    curves were compared with the log-rank test (p = 0.61, no significant difference). KM correctly handles censored time
    data (those whose event has not yet occurred / still alive); a plain average ignores these observations and is
    biased. In medicine it is the gold standard for overall survival and time-to-adverse-outcome analysis.

====================================================================================

#87  Cox Proportional Hazards
    file: 87_cox_death_hazard.xlsx
  >> SCENARIO (narration):
    We examine death risk with continuous and categorical predictors in a Cox model.
    For multiple predictors in censored time, Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death
    - Predictor(s): age
    - Predictor(s): stage
    - Predictor(s): new_treatment
    - Predictor(s): biomarker

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.595   age: HR = 1.013   p = 0.063 (borderline ns)   time: duration_month, event: death
    predictors: age, stage, new_treatment, sepsis biomarker

>> COMMENTARY (narration):
    We examined death risk with continuous and categorical predictors in a Cox model: the age effect is borderline
    (HR = 1.013, p = 0.063), concordance 0.60 (moderate discrimination). Cox regression gives the effect of several
    predictors on "event time" as hazard ratios (HR) in censored time data, making no assumption about the baseline
    hazard's shape. In medicine it is the gold standard for identifying the prognostic factors (age, stage, treatment,
    sepsis biomarker) that drive survival.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT)
    file: 88_aft_weibull_diyaliz.xlsx
  >> SCENARIO (narration):
    We model survival time assuming a Weibull distribution. For explicit time
    estimation, AFT is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death
    - Predictor(s): diabetes

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution: Weibull   Median survival = 40.30 months   time: duration_month, event: death, predictor: diabetes

>> COMMENTARY (narration):
    We modeled survival time by assuming a parametric (Weibull) distribution: median survival ~40.3 months. AFT
    (accelerated failure time) models, unlike Cox, choose an explicit distribution for the hazard shape and directly
    interpret how predictors "accelerate/decelerate" time. If the data fit the assumed distribution they are more
    powerful than Cox. In medicine they are preferred when explicit survival-time estimation and extrapolation are
    needed.

====================================================================================

#89  Competing Risks
    file: 89_competing_risks_3neden.xlsx
  >> SCENARIO (narration):
    We model a setting where a patient can meet several distinct death causes. For
    mutually exclusive events, competing risks is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death_cause -> olay tipi / event type
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 63   different death causes (death_cause) modeled as distinct event types

>> COMMENTARY (narration):
    We examined a setting where a patient can meet more than one distinct death cause (death_cause: disease-related,
    cardiac, other) in the competing-risks framework: 63 observations censored, the rest split into distinct event
    types. While standard survival treats all events alike, the competing-risks method accounts for the fact that "once
    one occurs the others no longer can" and gives a separate cumulative incidence for each death cause. In medicine it
    is necessary to correctly model a patient's mutually exclusive distinct death causes.

====================================================================================

#90  Time-Dependent Cox
    file: 90_tvcox_SpO2_death.xlsx
  >> SCENARIO (narration):
    We build a Cox model where the predictor (SpO2) changes over time. For a
    time-varying covariate, time-dependent Cox is appropriate.
  >> VARIABLE SELECTION:
    - Unit (id): patient_id
    - Start: start
    - Stop: end
    - Event: event
    - Predictor(s): SpO2_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    SpO2: HR = 1.096   p = 0.576 ns   (time-varying covariate)   id: patient_id, start/stop intervals

>> COMMENTARY (narration):
    We built a Cox model where the predictor (SpO2) CHANGES over time: for each patient the current SpO2 value was used
    via start-stop intervals; the effect here is non-significant (HR = 1.096, p = 0.58). Time-dependent Cox correctly
    handles non-constant covariates (changing SpO2, changing clinical status) -- it uses the predictor's current value
    at the event time. In medicine it is the right method for modeling the effect of time-varying clinical values
    (labs, treatment) on an event.

====================================================================================

#91  Survey Cox Regression
    file: 91_survey_phreg_death.xlsx
  >> SCENARIO (narration):
    We carry survival analysis into a complex survey design (weight+cluster). For
    design-faithful event time, survey_phreg is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death
    - Predictor(s): region (faktorize/factorized)
    - Weight: weight
    - Cluster: clinical_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.523   region_n: p = 0.445 ns   time: duration_month, event: death   weight: weight, cluster: clinical_id

>> COMMENTARY (narration):
    We carried survival analysis into a complex survey design (weight + cluster) with a Cox model: the region effect is
    non-significant (p = 0.45), concordance 0.52 (weak discrimination). survey_phreg extends the proportional-hazards
    model to a stratum/cluster/weight structure -- standard errors are corrected for the design. In medicine it is the
    design-faithful way to model event time (death) in representative clinical-registry/cohort data.

====================================================================================

#92  Interval-Censored Survival
    file: 92_interval_censored_nuks.xlsx
  >> SCENARIO (narration):
    We model data where the event is known only within an interval. For interval
    censoring, this is appropriate.
  >> VARIABLE SELECTION:
    - Lower bound: left_censor
    - Upper bound: survival_censor

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of events = 30   Median survival = 10.0   (lower bound: left_censor, upper bound: survival_censor)

>> COMMENTARY (narration):
    We modeled data where the exact event time is unknown and only known to have occurred within an INTERVAL: median
    survival 10 units, 30 events. Interval-censored methods are for cases where the event lies "somewhere between two
    observations" (between periodic checks); fixing the event to the interval's mid/end point biases results, while this
    method carries the uncertainty correctly. In medicine it is used to correctly model events occurring between
    periodic visits (recurrence/progression between two visits).

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty_cox_clinical.xlsx
  >> SCENARIO (narration):
    We model the center-specific hidden risk as frailty in clustered survival data.
    For shared hidden risk, frailty Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death
    - Predictor(s): age
    - Cluster: clinical_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.513   age: p = 0.530 ns   time: duration_month, event: death   cluster: clinical_id (frailty)

>> COMMENTARY (narration):
    In clustered survival data (patients within the same clinic/center) we modeled the center-specific unobserved risk as
    "frailty" (a random effect); the age effect is non-significant (p = 0.53). Frailty Cox accounts for the hidden risk
    shared by units in the same cluster -- solving the independence assumption that standard Cox violates. In medicine it
    is the right choice in multi-center studies for event data clustered within a center/clinic (shared hidden risk).

====================================================================================

#94  Time Series Analysis
    file: 94_ts_monthly_case.xlsx
  >> SCENARIO (narration):
    We examine a monthly case series (trend, season, stationarity). For
    time-dependent structure, time series analysis is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: monthly_case

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.998 (not stationary)   Trend: increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly case series (monthly_case): the ADF test shows it is not stationary (p = 0.998), with an
    increasing trend and seasonality. Time series analysis reveals the time-dependent structure (trend, season,
    autocorrelation) of observations; ordinary statistics are misleading because of this dependency. Stationarity is a
    prerequisite for models like ARIMA; if non-stationary, differencing is needed. In medicine/epidemiology it is the
    starting step for analyzing case/admission series.

====================================================================================

#95  STL Decomposition
    file: 95_stl_YBU_daily.xlsx
  >> SCENARIO (narration):
    We decompose the ICU occupancy series into trend, season and residual. For
    seasonal decomposition, STL is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: YBU_occupancy

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Season period = 7   series decomposed into trend + season + residual components

>> COMMENTARY (narration):
    We decomposed the ICU occupancy series (YBU_occupancy) into three components with STL: trend, the seasonal pattern
    (period = 7), and residual. STL visually and numerically separates a time series' long-term trend, recurring
    seasonal pattern and unexplained fluctuation; this answers "what is the underlying trend, and how much is seasonal?".
    In medicine/health management it is fundamental for de-seasonalizing occupancy/case series (epidemic season) to see
    the underlying trend and for anomaly detection.

====================================================================================

#96  ARIMA Forecast
    file: 96_arima_monthly_ddd.xlsx
  >> SCENARIO (narration):
    We model the monthly DDD (anesthetic consumption) series with ARIMA and produce a
    forecast. For series forecasting, ARIMA is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: monthly_ddd

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model selected   AIC = 2471.69   forecast + confidence interval

>> COMMENTARY (narration):
    We modeled the monthly DDD (defined daily dose / anesthetic consumption) series with ARIMA and produced a forecast
    (AIC = 2471.69 for model selection). ARIMA forecasts the future from the series' own past values (AR), trend (I -
    differencing) and past errors (MA); on a stationarized series it gives strong short-to-medium-term forecasts. In
    medicine/health management it is the most common classical method for anesthetic-consumption/case/demand forecasting; the
    confidence interval shows the forecast uncertainty.

====================================================================================

#97  Exponential Smoothing (ETS)
    file: 97_ets_monthly_emergency.xlsx
  >> SCENARIO (narration):
    We model the emergency-application series with Holt-Winters exponential
    smoothing. For a seasonal trended series, ETS is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: emergency_application

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method: Holt-Winters (seasonal additive, period = 12)   AIC = 1062.52

>> COMMENTARY (narration):
    We modeled the emergency-application series (emergency_application) with Holt-Winters exponential smoothing: it
    jointly estimates the level, trend and a 12-period seasonal component. ETS tracks the series' current level, trend
    and season by weighting recent observations more (exponentially decaying weights); for seasonal and trended series
    it is a practical alternative to ARIMA. In medicine/health management it gives fast, reliable results for emergency-
    application/load forecasting with regular seasonal patterns.

====================================================================================

#98  Mann-Kendall Trend
    file: 98_mann_kendall_survival_expectation.xlsx
  >> SCENARIO (narration):
    We test a significant trend in the life-expectancy series without distributional
    assumptions. For robust trend detection, Mann-Kendall is appropriate.
  >> VARIABLE SELECTION:
    - Value: survival_expectation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   significant trend detected   (trend magnitude via Sen's slope)   variable: survival_expectation

>> COMMENTARY (narration):
    We tested whether the life-expectancy series has a significant trend -- without assuming a distribution -- with
    Mann-Kendall: p < .001, a significant trend; Sen's slope gives the robust (outlier-resistant) magnitude of the
    trend. Mann-Kendall is a nonparametric trend test; it requires no normality and is resistant to outliers, hence
    common in health/epidemiology time series. In medicine it is used to robustly detect long-term trends in health
    indicators (rising/falling trend).

====================================================================================

#99  Anomaly Detection
    file: 99_anomali_lab_results.xlsx
  >> SCENARIO (narration):
    We detect unusual observations in multivariate laboratory data. For composite
    outlier detection, anomaly detection is appropriate.
  >> VARIABLE SELECTION:
    - Variables: WBC_K_uL
    - Variables: Hb_g_dL
    - Variables: creatinine
    - Variables: AST_U_L
    - Variables: procalcitonin_ng_mL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 17   variables: WBC_K_uL, Hb_g_dL, creatinine, AST_U_L, CRP_mg_L

>> COMMENTARY (narration):
    We automatically detected unusual observations in multivariate laboratory data: 17 patients were flagged as
    anomalies. Anomaly detection finds observations that deviate markedly from normal (erroneous record, critical/
    exceptional clinical picture) by evaluating several variables together; it catches "composite" outliers that
    univariate thresholds miss. In medicine it is used for laboratory quality control, erroneous-measurement detection
    and the early detection of critical patient profiles.

====================================================================================

#100  Variance Components
    file: 100_varcomp_clinical_doctor_patient_h2.xlsx
  >> SCENARIO (narration):
    We decompose treatment-response variability into nested levels (clinic/doctor).
    For hierarchical variability, variance components is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: treatment_response
    - Factor (categorical): clinical
    - Factor (categorical): doctor_id (nested)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Random factors: clinical + doctor_id (nested within clinical)   % contribution of each component + residual

>> COMMENTARY (narration):
    We decomposed treatment-response variability into nested levels -- clinic and doctor within clinic: the percent
    contribution of each level and the residual were reported. Variance components analysis answers "how much of the
    variability is between clinics, how much between doctors, how much within patient?". In medicine it is used in
    multi-center hierarchical structures (clinic>doctor>patient) to see where uncertainty concentrates and in sampling
    design.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old_treatment.xlsx
  >> SCENARIO (narration):
    We examine two treatment groups' response-score difference with a Bayesian
    t-test, expressing evidence as a Bayes factor. For an intuitive evidence ratio,
    this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: response_score
    - Grouping (categorical): treatment

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 2.81e+40   Cohen's d = 3.937   (overwhelming evidence for H1)   outcome: response_score, group: treatment

>> COMMENTARY (narration):
    We examined the response-score difference between two treatment groups with a Bayesian t-test: the Bayes factor is
    overwhelming (BF10 ~ 2.8e40), the effect very large (d = 3.94) -- the data support the "difference" hypothesis over
    "no difference" by astronomical odds. Unlike a p-value, the Bayes factor gives the RELATIVE evidence strength of two
    hypotheses and can distinguish "no evidence" from "no difference". In medicine it is preferred when one wants to
    express the evidential strength of a treatment decision as an intuitive ratio.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BMI_SpO2_BF10.xlsx
  >> SCENARIO (narration):
    We evaluate the relationship between BMI and SpO2 in a Bayesian framework. For
    evidential strength, Bayesian correlation is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: BMI
    - 2nd measure / group: SpO2_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.662 (strong)   BF10 = 4.77e+08   (overwhelming evidence for a relationship)

>> COMMENTARY (narration):
    We evaluated the relationship between two variables (BMI x SpO2) in a Bayesian framework: r = 0.66 and BF10 ~ 4.8e8,
    i.e. very strong evidence for a relationship. Bayesian correlation, instead of a classical p-value, presents the
    evidential strength of the relationship as a Bayes factor and the coefficient's posterior distribution. In medicine
    it is valuable for reporting not just whether the relationship between two clinical indicators is "significant" but
    how strongly it is "evidenced".

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_anesthetic_sex.xlsx
  >> SCENARIO (narration):
    We examine SpO2's difference across two factors (anesthetic, sex) with Bayesian
    ANOVA. For the evidential strength of factor effects, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: SpO2_pct
    - Factor (categorical): anesthetic
    - 2nd Factor: sex

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    anesthetic: BF10 = 2.59e+08 (very strong evidence)   outcome: SpO2, factors: anesthetic, sex

>> COMMENTARY (narration):
    We examined SpO2's difference across two factors (anesthetic, sex) with Bayesian ANOVA: for anesthetic, BF10 ~ 2.6e8, very
    strong evidence. Bayesian ANOVA compares the effects of factors and their interactions via Bayes factors; it ranks
    probabilistically which model (which effects) best explains the data. Unlike classical ANOVA's "reject/don't reject"
    decision, it quantifies the relative support among models. In medicine it is used to compare the evidential strength
    of factor (anesthetic) effects.

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_clinical_random.xlsx
  >> SCENARIO (narration):
    We analyze clinic-nested data with a Bayesian hierarchical model. For stable
    estimates in multi-center data, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: SpO2_pct
    - Cluster: clinical
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2_u (group) = 0.491 (46.5%)   sigma^2_eps (residual) = 0.566 (53.5%)   outcome: SpO2, group: clinical

>> COMMENTARY (narration):
    We analyzed clinic-nested data with a Bayesian hierarchical (multilevel) model: 46.5% of variability is
    between-clinic, 53.5% within-clinic (ICC ~ 0.47 -- high). This shows clinics make a large difference in SpO2
    outcomes. The Bayesian hierarchical model is the Bayesian version of LMM -- it estimates group effects with posterior
    distributions and balances small groups via "partial pooling". In medicine it is powerful for producing stable
    estimates even in small groups within multi-center (clinic/doctor/patient) data.

====================================================================================

#105  Spatial SAR
    file: 105_spatial_sar_COVID_case_ratio.xlsx
  >> SCENARIO (narration):
    When modeling province-level case ratio we handle spatial spillover with SAR.
    For neighborhood/contagion effects, SAR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: case_ratio
    - Predictor(s): socioeconomic
    - Predictor(s): age_mean
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho = 0.319   z = 3.95   p < .001 ***   Pseudo R^2 = 0.77   (N = 100, k-NN W, k = 5)
    outcome: case_ratio, predictors: socioeconomic, age_mean

>> COMMENTARY (narration):
    When modeling the province-level case ratio (case_ratio), we handled spatial spillover (the effect of neighboring
    provinces) with SAR: the spatial lag parameter is significant and positive (rho = 0.32, p < .001) -- a province's
    case ratio is related to its neighbors', a "cluster/spread" pattern. SAR incorporates spatial dependency into the
    model; if ignored, standard errors are biased. In medicine/epidemiology it is the right method for modeling the
    geographic spread of disease-rate indicators (neighborhood/contagion effect).

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_disease_residual_autocorr.xlsx
  >> SCENARIO (narration):
    We model spatial dependency in the error term. For unmeasured geographic
    factors, SEM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: case_ratio
    - Predictor(s): socioeconomic
    - Predictor(s): age_mean
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda = 0.630   z = 7.08   p < .001 ***   Pseudo R^2 = 0.72   (N = 100, k-NN W, k = 5)

>> COMMENTARY (narration):
    This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
    0.63, p < .001) -- the effect of geographic variables omitted from the model makes neighboring errors correlated.
    Unlike SAR, SEM attributes the spread to the error rather than the outcome. In medicine/epidemiology, when the source
    of spatial autocorrelation is unmeasured geographic factors (health access, environment), the correct specification
    is SEM; it is chosen by comparison with SAR.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local_risk_factor.xlsx
  >> SCENARIO (narration):
    Assuming the relationship is not constant in space, we estimate separate
    coefficients per province. For spatial heterogeneity, GWR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: case_ratio
    - Predictor(s): socioeconomic
    - Predictor(s): age_mean
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.811   local coefficients vary by location   outcome: case_ratio, predictors: socioeconomic, age_mean

>> COMMENTARY (narration):
    Assuming the relationship is NOT constant across space, we estimated SEPARATE coefficients for each province/location
    (R^2 = 0.81). Unlike "global" regression, GWR fits a separate model at each point with local overlap/weights; this
    answers "how does this variable's effect vary by region?" and maps spatial heterogeneity. In medicine/epidemiology
    it is powerful where the socioeconomic-disease relationship differs by region.

====================================================================================

#108  Nested Mixed Model
    file: 108_nested_lmm_clinical_doctor_visit_SpO2.xlsx
  >> SCENARIO (narration):
    We model SpO2 in a nested design (clinic>doctor>visit). For nested hierarchies,
    nested LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: SpO2_pct
    - Factor (categorical): visit
    - Factor (categorical): clinical
    - Factor (categorical): doctor_no

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    clinical (P) effect: F = 10.37, p < .001 ***   doctor_no (within clinic, F): F = 7.57, p < .001 ***
    visit (R) effect: F = 14.20, p < .001 ***   outcome: SpO2

>> COMMENTARY (narration):
    We modeled SpO2 in a nested design (clinic > doctor > visit): the clinic level (F = 10.37, p < .001), the doctor
    level (F = 7.57, p < .001) and the visit/replication level (F = 14.20, p < .001) all make significant contributions.
    Nested LMM correctly handles hierarchies where sub-units are nested within super-units (each doctor belongs to only
    one clinic); by partitioning variance into levels it shows each layer's share. In medicine it is used to correctly
    separate effects in clinic>doctor>patient hierarchies.

====================================================================================

#109  Crossed Mixed Model
    file: 109_crossed_lmm_anesthetic_clinical.xlsx
  >> SCENARIO (narration):
    We model a design where two random factors (anesthetic x clinic) are crossed. For two
    independent classification axes, crossed LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: SpO2_pct
    - Factor (categorical): anesthetic
    - 2nd Factor: clinical

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    anesthetic (A): F = 2.24, p = 0.125 ns   clinical (B): F = 4.31, p = 0.028 *   A x B: F = 13.00, p < .001 ***   outcome: SpO2

>> COMMENTARY (narration):
    We modeled a design where two random factors are crossed rather than nested (each anesthetic is given in each clinic): the
    anesthetic main effect is non-significant (p = 0.13), the clinic main effect is significant (p = 0.028), and the A x B
    interaction is very strong (F = 13.00, p < .001) -- so an anesthetic's effect DEPENDS on which clinic it is given in.
    Crossed LMM, unlike nested, handles two independent grouping axes (anesthetic x clinic, each anesthetic in each clinic) at once.
    In medicine it is the right choice for jointly analyzing the effects and interaction of two independent
    classification axes.

====================================================================================

#110  Kernel Density (KDE) Map
    file: 110_KDE_hastane_density_haritasi.xlsx
  >> SCENARIO (narration):
    We produce a continuous density surface from unit locations. For a
    case-concentration map, KDE is appropriate.
  >> VARIABLE SELECTION:
    - Value: case_ratio
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 95 units   case-ratio-weighted spatial density surface (continuous heat map)

>> COMMENTARY (narration):
    Using unit locations (n = 95) we produced a continuous density surface (a KDE heat map): the point distribution was
    turned into a smooth "density" surface. KDE answers "where is it dense?" from scattered point data as a continuous
    map; it shows the trend rather than individual points. In medicine/epidemiology it is the core spatial tool for
    visually mapping case/disease concentrations and identifying empty/dense zones.

====================================================================================

#111  Hexbin Density Map
    file: 111_Hexbin_case_distribution.xlsx
  >> SCENARIO (narration):
    We aggregate unit distribution into hexagonal cells to show density. For the
    over-plotting problem, hexbin is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 150 units   spatial density via aggregation into hexagonal cells

>> COMMENTARY (narration):
    We colored unit distribution (n = 150) by counting within hexagonal cells. Hexbin aggregates many overlapping points
    (the over-plotting problem) into a regular hexagonal grid to show density clearly; compared with squares it carries
    less directional bias. In medicine/epidemiology it is used to turn dense point clouds (case locations) into a
    readable density map and to compare spatial concentrations.

====================================================================================

#112  Moran's I
    file: 112_Morans_I_COVID_spatial_autocorrelation.xlsx
  >> SCENARIO (narration):
    We test whether the case ratio is distributed randomly or in clusters across
    space. For spatial autocorrelation, Moran's I is appropriate.
  >> VARIABLE SELECTION:
    - Value: case_ratio
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.684   z = 12.01   p < .001 ***   (k-NN neighbors, k = 6)   variable: case_ratio

>> COMMENTARY (narration):
    We tested whether the case ratio is distributed randomly or in clusters across space with Moran's I: I = 0.68,
    z = 12.01, p < .001 -- very strong positive spatial autocorrelation, i.e. similar case ratios cluster
    geographically. Moran's I quantifies "Tobler's first law" (near things are similar). In medicine/epidemiology it is
    the first test for detecting the geographic clustering of disease rates (outbreak patterns) and for deciding whether
    a spatial model is needed.

====================================================================================

#113  Getis-Ord Gi*
    file: 113_Getis_Ord_disease_hotspot.xlsx
  >> SCENARIO (narration):
    We map locally where the case ratio clusters high/low. For hot/cold spots,
    Getis-Ord is appropriate.
  >> VARIABLE SELECTION:
    - Value: case_ratio
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Hot spots = 33   Cold spots = 30   n = 68   (k-NN k = 8, binary weights)   variable: case_ratio

>> COMMENTARY (narration):
    We mapped locally WHERE the case ratio clusters high/low with Getis-Ord Gi*: 33 significant hot spots (high-case-rate
    clusters) and 30 cold spots (low-rate clusters). While Moran's I states the overall clustering, Gi* shows its
    location -- for each point it tests "is the surrounding area high or low?". In medicine/public health it is used to
    pinpoint high-risk zones (hot spots) for targeted intervention/resource allocation.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_case_kumeleri.xlsx
  >> SCENARIO (narration):
    We cluster unit locations with density-based DBSCAN. For spatial clusters and
    outlier locations, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Clusters = 4   Noise (outliers) = 11   n = 115   (location: lat, lon)

>> COMMENTARY (narration):
    We clustered unit locations (n = 115) with density-based DBSCAN: 4 natural geographic clusters and 11 "noise"
    (scattered, cluster-less) points. DBSCAN does not require the cluster count in advance, finds clusters of any shape,
    and separates sparse points as outliers -- ideal for spatial clustering. In medicine/epidemiology it is used to
    detect region-level natural case concentrations (outbreak foci) and to isolate isolated locations; it completes our
    descriptive spatial analysis series.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_pain_VAS.xlsx
  >> SCENARIO (narration):
    We follow 40 patients measured at three time levels (1h, 6h, 24h); each belongs to one of two anesthesia_technique groups (general / regional). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on postoperative pain (VAS).
  >> VARIABLE SELECTION:
    - Dependent variable: pain_VAS
    - Subject ID: patient_id
    - Between-subjects factor: anesthesia_technique
    - Within-subjects factor: time

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (anesthesia_technique): F(1,38) = 26.37  p < .001  np2 = 0.410
    Within (time):  F(2,76) = 148.05  p < .001  np2 = 0.796
    Interaction:         F(2,76) = 7.89  p < .001  np2 = 0.172
    Mauchly W = 0.886  p = 0.100   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a anesthesia_technique x time mixed design we analyzed postoperative pain (VAS) for 40 patients (120 observations). The interaction is significant (F(2,76) = 7.89, p < .001, np2 = 0.172) *** -- the two groups' change across time differs in magnitude. The between-subjects main effect (general vs regional) is F = 26.37, p < .001; the within-subjects main effect (1h/6h/24h) is F = 148.05, p < .001. Mauchly's test p = 0.100, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each time level. In anesthesia and critical care, the mixed design is the standard analysis for comparing anesthesia techniques on postoperative recovery over time.

====================================================================================
