====================================================================================
MerQur - Enfeksiyon_Hastaliklari - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Enfeksiyon_Hastaliklari/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_patient_vital.xlsx
  >> SCENARIO (narration):
    Suppose we have just admitted a cohort of patients to the infectious disease
    ward and we want a clear snapshot of who they are before any modelling.
    Across 280 patients we recorded age, sex, height, weight, BMI, systolic and
    diastolic blood pressure, and admission lactate. Descriptive statistics let
    us summarize central tendency and spread for the continuous markers and
    frequencies for sex, so we can describe the clinical profile of the cohort
    and spot anything unusual, like an unexpectedly high mean lactate suggesting
    tissue hypoperfusion. This is the natural first step in any infection study
    before we run inferential tests.
  >> VARIABLE SELECTION:
    - Variables (numeric): age, height_cm, weight_kg, BMI, systolic_bp, diastolic_bp, lactate_mg_dL
    - Grouping/categorical variable: sex (female, male)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 patients
    age      : mean = 51.9        |   BMI : mean = 27.3        |   systolic BP : mean = 134.0 mmHg
    lactate  : mean = 3.6 mg/dL

>> COMMENTARY (narration):
    We first drew the general clinical profile of our patient sample: 280 patients, mean age 52, BMI 27.3 (mildly
    overweight), systolic blood pressure 134 mmHg (mildly high), lactate 3.6 (elevated -- a sign of tissue
    hypoperfusion/infection). This descriptive table sets the stage for all the analyses that follow -- treatment
    comparisons, biomarker relationships, survival models. In clinical research, before any inferential test,
    summarising the patient population's basic demographic and vital characteristics is essential both to check data
    quality and to place the findings in the right context.

====================================================================================

#2  Normality Tests
    file: 02_normality_BMI_bp_procalcitonin.xlsx
  >> SCENARIO (narration):
    Before we choose between parametric and non-parametric tests for our septic
    patients, we need to know whether our key measurements are normally
    distributed. In 200 patients we measured BMI, systolic blood pressure, and
    procalcitonin, a biomarker that is famously right-skewed in infection.
    Normality tests, together with the accompanying Q-Q plots, tell us whether
    each variable can be analyzed with t-tests and ANOVA or whether we should
    consider transformations or rank-based alternatives. We expect procalcitonin
    to deviate from normality, and this analysis confirms it formally.
  >> VARIABLE SELECTION:
    - Variables to test for normality: BMI, systolic_bp, procalcitonin_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BMI               : Shapiro-Wilk = 0.996  p = 0.881   KS p = 0.917   Normal
    systolic_bp       : Shapiro-Wilk = 0.991  p = 0.275   KS p = 0.252   Normal
    procalcitonin_pct : Shapiro-Wilk = 0.939  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three clinical variables -- BMI, systolic blood pressure and procalcitonin -- are normally
    distributed. The result splits instructively: BMI and systolic blood pressure are normal (p > 0.05), but
    procalcitonin deviates significantly from normality (p < .001). This is typical: physiological measures (BMI,
    blood pressure) are usually near-normal, while most infection biomarkers (CRP, procalcitonin, ferritin) are
    right-skewed; low in most patients, very high in severe infection. The practical takeaway: we can safely use
    parametric tests (t-test, ANOVA) on BMI and blood pressure; for procalcitonin, nonparametric methods (Mann-Whitney,
    Kruskal-Wallis) or a log-transform are more appropriate. The normality check is a critical pre-step that decides,
    for each variable separately, which test family is appropriate.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_systolic_bp.xlsx
  >> SCENARIO (narration):
    Suppose clinical guidelines say the target mean systolic blood pressure for
    stabilized infection patients should be 120 mmHg. We measured systolic blood
    pressure in 100 patients on our ward and want to know whether our cohort's
    mean differs significantly from that reference value. A one-sample t-test is
    appropriate because we have a single continuous outcome compared against a
    known fixed standard, letting us judge whether our patients are
    systematically above or below the clinical target.
  >> VARIABLE SELECTION:
    - Test variable: systolic_bp
    - Test value (population mean): 120

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = 5.8605   p < .001 ***   Cohen d = 0.586 (Moderate)
    Mean systolic BP = 128.25   (test mu = 120, normal threshold)   H0 REJECTED

>> COMMENTARY (narration):
    We compared a patient group's mean systolic blood pressure against the 120 mmHg normal reference threshold. The
    result: mean 128.3, significantly above the threshold -- t(99) = 5.86, p < .001, moderate effect size d = 0.59.
    So this group's blood pressure exceeds the normal threshold by more than chance can explain. The one-sample
    t-test is the right way to compare a clinical measurement against a known reference/guideline value (normal limit,
    target value) -- a very common question in medical research. The effect size shows the difference is clinically
    meaningful too.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent Samples t-Test
    file: 04_independent_t_antibiotic_procalcitonin.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether a new antibiotic lowers procalcitonin more
    than placebo. In this trial 115 septic patients were randomized to either
    the antibiotic or placebo arm, and we recorded the change in procalcitonin
    from baseline. We use an independent samples t-test because the outcome,
    procalcitonin change, is continuous and the predictor is a single two-level
    group, so we are comparing the mean reduction between two independent sets
    of patients.
  >> VARIABLE SELECTION:
    - Group (2 levels): group (antibiotic, placebo)
    - Dependent variable: procalcitonin_change

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = -10.816   p < .001 ***   Cohen d = -2.019 (Large)
    Procalcitonin change differs significantly by group.   H0 REJECTED

>> COMMENTARY (narration):
    We compared procalcitonin change between two treatment groups (new vs standard antibiotic) with an independent
    t-test. The difference is highly significant and large (t(113) = -10.82, p < .001, d = -2.02): one group's
    procalcitonin fell markedly more -- the infection response was better controlled. Such a large effect points to a
    clinically very important treatment advantage. The independent t-test is the basic method for comparing a
    continuous outcome (biomarker change) between two independent treatment/patient groups.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired t-Test
    file: 05_paired_t_bp_treatment.xlsx
  >> SCENARIO (narration):
    Suppose we start 50 patients on a wide-spectrum antibiotic and want to know
    whether their blood pressure changes over the treatment course. For each
    patient we recorded systolic blood pressure before treatment and again after
    several weeks on therapy. Because both measurements come from the same
    patient, the observations are paired, so a paired t-test is the right choice
    to test whether the mean within-patient change in systolic pressure is
    significantly different from zero.
  >> VARIABLE SELECTION:
    - Paired variable 1 (before): systolic_before
    - Paired variable 2 (after): systolic_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Paired t(49) = 18.4314   p < .001 ***   Cohen d_z = 2.607 (Large)
    Mean difference (before − after systolic) = 19.98 mmHg   H0 REJECTED

>> COMMENTARY (narration):
    We measured a treatment's effect on blood pressure by comparing the same patients' systolic values before and
    after with a paired t-test. The result is striking: blood pressure fell by 20 mmHg on average, t(49) = 18.43,
    p < .001, with an enormous effect size (d_z = 2.61). The treatment lowered blood pressure very strongly. Because
    we measured the same patients twice, the paired test is correct -- it removes between-patient variability and
    focuses only on the patient's own change. It is the gold-standard way to evaluate before/after treatment in
    medical research.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_antibiotic_ferritin.xlsx
  >> SCENARIO (narration):
    Suppose we want to compare three antibiotics against placebo for their
    effect on inflammation. We randomized 140 patients into four arms, placebo,
    antibiotic-a, antibiotic-b, and antibiotic-c, and measured the change in
    ferritin, an acute-phase marker. A one-way ANOVA is appropriate because we
    have one continuous outcome, ferritin change, and a single categorical
    factor with more than two levels, allowing us to test whether at least one
    antibiotic differs from the others before running post-hoc comparisons.
  >> VARIABLE SELECTION:
    - Factor (4 levels): antibiotic_group (placebo, antibiotic-a, antibiotic-b, antibiotic-c)
    - Dependent variable: ferritin_change

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 136) = 57.6701   p < .001 ***   η² = 0.5599
    DECISION: H0 REJECTED (4 antibiotic groups)

>> COMMENTARY (narration):
    We compared the effect of four antibiotic groups on ferritin change with one-way ANOVA. The result is highly
    significant (F(3,136) = 57.67, p < .001), with effect size η² = 0.56 -- more than half of the ferritin change is
    explained by antibiotic group, a very strong effect. ANOVA says at least one group differs from the others; a
    post-hoc test is needed to see which antibiotics differ. "Which treatment/drug is more effective" is a
    fundamental question in medical research; ANOVA is the standard way to answer it across multiple groups.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova_antibiotic_sex.xlsx
  >> SCENARIO (narration):
    Suppose we suspect that an antibiotic's effect on blood pressure may differ
    between men and women. In 150 patients we crossed three antibiotics,
    antibiotic-a, antibiotic-b, and antibiotic-c, with sex, and measured the
    change in blood pressure. A two-way ANOVA lets us test the main effect of
    antibiotic, the main effect of sex, and crucially their interaction, so we
    can see whether the antibiotic response depends on the patient's sex.
  >> VARIABLE SELECTION:
    - Factor 1 (3 levels): antibiotic (antibiotic-a, antibiotic-b, antibiotic-c)
    - Factor 2 (2 levels): sex (female, male)
    - Dependent variable: bp_change

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    antibiotic (main) : F(2) = 29.95   p < .001 ***   η²p = 0.29
    sex (main)        : F(1) =  9.04   p = 0.003 **    η²p = 0.06

>> COMMENTARY (narration):
    We examined blood-pressure change with two factors -- antibiotic type and sex -- together. The antibiotic effect
    is strong (F = 29.95, η²p = 0.29), while the sex effect is significant but small (F = 9.04, p = 0.003, η²p = 0.06).
    So the main determinant of the outcome is treatment type; sex contributes only modestly. Two-way ANOVA's strength
    is seeing both factors' effects -- separately and jointly (interaction) -- at once. This answers questions like
    "is the treatment effect similar in both sexes" and informs the basic questions of personalised medicine.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated Measures ANOVA
    file: 08_repeated_anova_visit_procalcitonin.xlsx
  >> SCENARIO (narration):
    Suppose we are tracking how procalcitonin falls as patients respond to
    treatment. We followed 80 patients and measured procalcitonin at four
    scheduled visits. Because the same patients are measured repeatedly over
    time, the measurements are correlated within each patient, so a repeated
    measures ANOVA is appropriate to test whether mean procalcitonin changes
    significantly across the four visits while accounting for the within-subject
    design.
  >> VARIABLE SELECTION:
    - Repeated measures (within-subject levels): visit_1_procalcitonin, visit_2_procalcitonin, visit_3_procalcitonin, visit_4_procalcitonin

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 237) = 16.0082   p < .001 ***   η²p = 0.1685   (n = 80)
    DECISION: H0 REJECTED

>> COMMENTARY (narration):
    We compared the same 80 patients' procalcitonin across four consecutive visits with repeated-measures ANOVA.
    Because the same patients are measured repeatedly, observations are dependent; RM-ANOVA accounts for this
    within-subject correlation. The result is significant (F(3,237) = 16.01, p < .001, η²p = 0.17): procalcitonin
    changes significantly across visits -- typically a decreasing trend with treatment. RM-ANOVA is the correct
    design for longitudinal treatment-monitoring studies where the same patients are tracked over time, and is far
    more powerful than treating each visit as independent.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_treatment_3vital.xlsx
  >> SCENARIO (narration):
    Suppose a treatment may affect several hemodynamic outcomes at once rather
    than just one. In 120 patients we compared placebo against treatment-A and
    treatment-B and recorded three correlated vital signs, systolic pressure,
    diastolic pressure, and pulse. MANOVA is the right tool here because we have
    multiple continuous dependent variables and a single categorical factor,
    letting us test whether the treatment groups differ on the combined
    hemodynamic profile while controlling the overall error rate.
  >> VARIABLE SELECTION:
    - Factor (3 levels): treatment (placebo, treatment-A, treatment-B)
    - Dependent variables (multivariate): systolic, diastolic, pulse

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.346 -> F(6, 230) = 26.83   p < .001 ***   (n = 120)
    Dependent: systolic, diastolic, pulse   |   Factor: treatment

>> COMMENTARY (narration):
    Here we have three correlated vital parameters -- systolic, diastolic blood pressure and pulse -- and one
    treatment factor. Running three separate ANOVAs would inflate the error rate and ignore the correlations among
    parameters. MANOVA tests all three jointly: Wilks' Lambda F = 26.83, p < .001, so the treatment strongly shifts
    the multivariate vital profile. MANOVA is the right approach when a treatment affects several related vital
    outcomes at once -- it controls the overall error rate. After a significant MANOVA, follow-up univariate tests
    show which parameter drives the effect.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova_physical_treatment_VAS.xlsx
  >> SCENARIO (narration):
    Suppose we are comparing rehabilitation strategies for post-infection pain,
    but we know baseline pain strongly predicts the outcome. We assigned 105
    patients to control, mobilization, or mobilization plus manual therapy and
    recorded pain on a VAS scale before and after the program. ANCOVA is
    appropriate because it tests whether post-treatment pain differs between
    groups while statistically adjusting for baseline pain as a covariate,
    giving us a fairer comparison of the treatment effect.
  >> VARIABLE SELECTION:
    - Factor (3 levels): treatment (control, mobilization, mobilization+manual)
    - Dependent variable: pain_post_VAS
    - Covariate: pain_baseline_VAS

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    treatment (factor): η²p = 0.712 (very large)   p < .001 ***
    Covariate: pain_baseline_VAS (baseline pain)   |   DV: pain_post_VAS   H0 REJECTED

>> COMMENTARY (narration):
    When evaluating a pain treatment's effect, it is essential to control for patients' baseline pain levels --
    because if the groups differ at baseline, the post-treatment pain difference is misleading. ANCOVA takes the
    baseline VAS as a covariate and holds it statistically constant. The result: even after adjusting for baseline
    pain, the treatment groups differ very significantly in post-treatment pain (η²p = 0.71, very large effect). So
    the observed difference is not a by-product of baseline differences; it is the genuine treatment effect. ANCOVA
    is the standard way to fairly control baseline differences in clinical intervention studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci_lactate.xlsx
  >> SCENARIO (narration):
    Suppose we want to estimate the average blood lactate level in critically
    ill infectious-disease patients, but our sample is small and lactate values
    are notoriously skewed. Here we have just 30 patients, each with a measured
    lactate_mmol_L, and we don't want to assume the data are normally
    distributed. So instead of relying on a classic t-based confidence interval,
    we use a bootstrap confidence interval, which resamples the observed lactate
    values thousands of times to build an empirical interval for the mean. This
    is ideal because it makes no parametric assumption and gives us a robust
    range for the true mean lactate even with only a handful of patients.
  >> VARIABLE SELECTION:
    - Variable / Statistic of interest: lactate_mmol_L

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean lactate = 4.07 mmol/L
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    We derived a confidence interval for the mean of lactate by bootstrapping: the observed mean is 4.07 mmol/L
    (elevated -- a sign of severe infection/sepsis). Clinical biomarkers like lactate are right-skewed, and classic
    normal-assuming formulas can be unreliable for such data. Bootstrap estimates the interval without distributional
    assumptions, through thousands of resamples. It is a robust way to express the uncertainty of a mean in skewed
    clinical data (lactate, CRP, length of stay).

====================================================================================

#12  Permutation Test
    file: 12_permutation_CRP_case_control.xlsx
  >> SCENARIO (narration):
    Imagine we suspect that CRP, a key inflammatory marker, is elevated in
    confirmed infection cases compared with culture-negative controls. We have
    53 patients split into two groups: those with confirmed infection labelled
    'patient' and the 'culture_negative' controls, and for each we recorded
    CRP_mg_L. With such a modest, possibly skewed sample, we don't want to trust
    the assumptions of a parametric test, so we run a permutation test. By
    repeatedly shuffling the group labels and recomputing the difference in mean
    CRP, we get an exact, assumption-free p-value for whether infection cases
    really have higher CRP.
  >> VARIABLE SELECTION:
    - Group (2 levels): group (culture_negative, patient)
    - Test variable: CRP_mg_L

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = -3.47   p = 0.0011 **   H0 REJECTED
    CRP differs significantly between case and control groups

>> COMMENTARY (narration):
    We tested the CRP difference between case (patient) and control groups with a permutation test. Permutation
    shuffles the group labels thousands of times to measure whether the observed difference could arise by chance --
    requiring no distributional assumption. The result is significant (p = 0.001): the cases' and controls' CRP
    differs more than chance allows. For small samples or skewed biomarkers like CRP where normality is doubtful,
    permutation is a robust, assumption-free alternative to the t-test. It is reliable in case-control studies.

====================================================================================

#13  Multiple Comparison Corrections
    file: 13_multiple_comparison_antikoagulan.xlsx
  >> SCENARIO (narration):
    Suppose we're studying anticoagulation control and want to compare the INR
    stability ratio across nine different antibiotic regimens, since several
    antibiotics are known to interact with warfarin. We have 180 patients, each
    on one of nine drugs (a reference warfarin-only group plus antibiotic-1
    through antibiotic-8), and the outcome is INR_stability_ratio. Running every
    pairwise comparison inflates the false-positive rate dramatically, so before
    drawing conclusions we apply multiple comparison corrections such as
    Bonferroni or Benjamini-Hochberg. This keeps the family-wise error under
    control and tells us which specific drug pairs truly differ in INR
    stability.
  >> VARIABLE SELECTION:
    - Group / Factor (9 levels): antibiotic (Varfarin, antibiotic-1, antibiotic-2, antibiotic-3, antibiotic-4, antibiotic-5, antibiotic-6, antibiotic-7, antibiotic-8)
    - Dependent variable: INR_stability_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Antibiotic groups vs reference: TOTAL SIGNIFICANT = 4/8 (in every correction method)

>> COMMENTARY (narration):
    When we compare different antibiotic groups pairwise on INR stability ratio, we run many tests -- each carrying a
    false-positive risk. Multiple-comparison correction controls that risk. Here four of eight comparisons stayed
    significant under every correction method (Bonferroni, Holm, FDR) -- 4/8. So some groups genuinely differ from the
    reference, while others lose significance after correction. Skipping this step risks reporting spurious
    differences; in drug/treatment comparisons the correct correction is essential for clinical reliability.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney_diabetes_survival_quality.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether survival-related quality of life differs
    between patients with type-1 and type-2 diabetes who are battling infection.
    We have 85 patients, classified by diabetes_type as type-1 or type-2, and
    for each we have a survival_quality score. Quality-of-life scores are
    ordinal and clearly non-normal, so a t-test would be inappropriate. Instead
    we use the Mann-Whitney U test, which compares the distributions of
    survival_quality between the two independent diabetes groups using ranks
    rather than means.
  >> VARIABLE SELECTION:
    - Group (2 levels): diabetes_type (type-1, type-2)
    - Test variable: survival_quality

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Mann-Whitney U = 886.50   p = 0.907 ns   r = .02
    Quality of life does NOT differ significantly by diabetes type.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether quality of life differs by diabetes type with the nonparametric Mann-Whitney. This time the
    result is non-significant (U = 886.5, p = 0.907, r = .02): there is no difference in quality of life between the
    two diabetes types. Non-significant results are valuable too -- always expecting a difference is a mistake. Here
    diabetes type is not a factor that distinguishes quality of life; it is probably determined by other factors
    (complications, social support). For ordinal/skewed quality-of-life scores, Mann-Whitney is more reliable than
    the t-test.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank Test
    file: 15_wilcoxon_treatment_VAS.xlsx
  >> SCENARIO (narration):
    Imagine we're evaluating whether a course of antibiotic treatment reduces
    patients' pain over time. We followed 35 infectious-disease patients and
    recorded their pain on a visual analogue scale before treatment, VAS_before,
    and again after treatment, VAS_post. These are paired measurements on the
    same patients, and VAS scores are not normally distributed, so a paired
    t-test isn't ideal. We therefore use the Wilcoxon signed-rank test, which
    examines the ranked differences within each patient to test whether pain
    dropped significantly after treatment.
  >> VARIABLE SELECTION:
    - Pair measurement 1 (before): VAS_before
    - Pair measurement 2 (after): VAS_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilcoxon W = 0.00   p < .001 ***   r = 0.888   n (nonzero) = 31
    VAS pain score changed significantly before vs after treatment.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same patients' VAS pain scores before and after a treatment with the paired Wilcoxon test. The
    result is very strong: W = 0, p < .001, effect size r = 0.89. Pain decreased in the same direction and by a large
    amount in nearly all patients -- the treatment had a clear effect. For paired measures and ordinal/skewed pain
    scores, Wilcoxon is the right choice. It is a robust, assumption-free way to measure pain/symptom severity before
    and after in clinical studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_antibiotic_duration.xlsx
  >> SCENARIO (narration):
    Suppose we want to compare how long antibiotic therapy lasts across three
    different drugs used for the same infection. We have 45 patients, each
    treated with one of three antibiotics: Amoksisilin, Sefuroksim, or
    Levofloksasin, and we recorded treatment_day, the number of days on therapy.
    With a small sample and skewed duration data across more than two groups,
    ANOVA assumptions are doubtful, so we use the Kruskal-Wallis test. It
    compares the rank distributions of treatment duration across all three
    antibiotic groups at once.
  >> VARIABLE SELECTION:
    - Group / Factor (3 levels): antibiotic (Amoksisilin, Sefuroksim, Levofloksasin)
    - Dependent variable: treatment_day

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 3.3487   p = 0.187 ns   eta-squared_H = 0.03
    Treatment duration does NOT differ significantly by antibiotic.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether treatment duration (days) varies by antibiotic type with Kruskal-Wallis -- the nonparametric
    ANOVA for more than two groups. The result is non-significant (H(2) = 3.35, p = 0.187): there is no difference in
    treatment duration among antibiotics. A non-significant result is information too -- the different antibiotics act
    over similar durations, so none is superior on duration. Because count/ordinal duration data may not be normally
    distributed, Kruskal-Wallis is more appropriate than classic ANOVA. Negative findings help avoid unnecessary
    treatment changes.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman_radiologist_method.xlsx
  >> SCENARIO (narration):
    Imagine we're validating four imaging-based diagnostic methods for detecting
    infection-related lesions. Thirty radiologists each evaluated all four
    methods, giving us repeated accuracy measurements per rater in the columns
    method_A_accuracy, method_B_accuracy, method_C_accuracy, and
    method_D_accuracy. Because the same radiologists rated every method, the
    measurements are related, and accuracy scores are not normally distributed.
    We therefore use the Friedman test, the non-parametric counterpart to
    repeated-measures ANOVA, to test whether the four methods differ in
    diagnostic accuracy.
  >> VARIABLE SELECTION:
    - Repeated measures (related conditions): method_A_accuracy, method_B_accuracy, method_C_accuracy, method_D_accuracy

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Friedman chi-square(3) = 32.08   p < .001 ***   Kendall W = 0.3564   n = 30
    The same radiologists' accuracy across 4 methods differs significantly.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the diagnostic accuracy the same radiologists achieved with four different imaging methods using the
    Friedman test -- the nonparametric counterpart of repeated-measures ANOVA. The result is highly significant
    (chi-square(3) = 32.08, p < .001), with Kendall W = 0.36 indicating moderate consistency of the ranking across
    methods. So the four methods make a clear difference in diagnostic accuracy. Because the same radiologists used
    all methods, observations are dependent; Friedman accounts for this correctly. It is the right test for medical
    studies comparing diagnostic methods on the same raters.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_surgery_achievement.xlsx
  >> SCENARIO (narration):
    Suppose a new surgical approach for source control in infected patients is
    claimed to succeed in 80 percent of cases, and we want to test that claim.
    We have 800 patients with a binary 'successful' outcome coded 1 for success
    and 0 for failure. Since the outcome is a simple yes-or-no event with a
    fixed hypothesized probability, the natural choice is a binomial test. It
    compares our observed proportion of successful surgeries against the
    hypothesized success rate to see whether the true success probability really
    differs from the stated benchmark.
  >> VARIABLE SELECTION:
    - Outcome (binary success): successful
    - Test proportion: 0.80

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed success proportion = 0.7375   test proportion = 0.50
    p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether a surgical procedure's success rate equals chance (50%) with the binomial test. The observed
    proportion is 0.74 -- three-quarters of procedures succeeded. This differs significantly from 50% (p < .001). The
    binomial test is the simplest way to compare a single binary proportion (success/failure, present/absent) against
    a known reference. In medical research it is a practical tool for comparing surgical/treatment success rates
    against a target or literature value.

====================================================================================

#19  Sign Test
    file: 19_sign_test_diet_weight.xlsx
  >> SCENARIO (narration):
    Imagine we want to know whether a nutritional intervention changes body
    weight in infectious-disease patients recovering from prolonged illness. We
    measured 28 patients' weight_before_kg and their weight_post_kg after the
    diet program. We're not willing to assume the weight changes are symmetric
    or normally distributed, and we only care about the direction of change for
    each patient. The sign test is the simplest choice: it just counts how many
    patients gained versus lost weight and tests whether increases and decreases
    are equally likely.
  >> VARIABLE SELECTION:
    - Pair measurement 1 (before): weight_before_kg
    - Pair measurement 2 (after): weight_post_kg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive (before > after) = 26   (weight mostly DECREASED)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same patients' weight before and after a diet program with the paired sign test. Weight decreased
    in the majority of patients (26 had a higher before-weight), p < .001. So the diet produced weight loss. The sign
    test looks only at the direction of change (up/down), not its magnitude; it is a robust, assumption-free test.
    When the direction of change (how many patients lost/gained) matters, or when the distribution is very skewed, it
    is a coarse but sturdy before/after indicator.

====================================================================================

#20  Runs Test
    file: 20_runs_test_YBU_occupancy.xlsx
  >> SCENARIO (narration):
    Suppose we're monitoring daily intensive-care bed occupancy during an
    outbreak and want to know whether the sequence of high and low occupancy
    days is random or shows clustering, such as sustained surges. We have 60
    consecutive days of bed_occupancy values recorded in day order. To detect
    non-randomness in this ordered sequence, we use the runs test. It
    dichotomizes occupancy around the median and counts the runs of consecutive
    above- and below-median days, testing whether there are too few or too many
    runs to be explained by chance.
  >> VARIABLE SELECTION:
    - Sequence variable (in time order): bed_occupancy
    - Order variable: day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 24   Z = -1.82   (borderline; accepted as random)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the daily ICU bed-occupancy series is randomly distributed -- whether successive days cluster
    or show a pattern -- with the Wald-Wolfowitz runs test. With Z = -1.82 the result is borderline but does not
    cross the significance threshold (usually |Z| > 1.96); the series is accepted as random. This matters
    methodologically: if a series is non-random (autocorrelated), many classic tests become invalid. The borderline Z
    suggests a slight clustering tendency but no strong pattern. The runs test is ideal for this pre-check on
    time-ordered data.

====================================================================================

#21  Chi-Square Test of Independence
    file: 21_chisquare_smoking_copd.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether smoking history is associated with chronic
    obstructive pulmonary disease in our infectious-disease cohort. We have
    records for 670 patients, where smoking_status is recorded in three levels -
    active smoking, old smoking, and none non_drinker - and copd is coded as
    present or absent. The clinical question is simple: is the rate of COPD
    different across smoking categories, or are the two variables independent?
    Because both variables are categorical and we are comparing observed counts
    against what we would expect under independence, the Chi-Square Test of
    Independence is the right tool here.
  >> VARIABLE SELECTION:
    - Row variable: smoking_status (active smoking, old smoking, none non_drinker)
    - Column variable: copd (present, absent)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(2) = 130.54   p < .001 ***   Cramer V = 0.44 (Strong)
    Smoking status is associated with COPD.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether smoking status is associated with COPD (chronic obstructive pulmonary disease) using
    chi-square. The result is very strong: chi-square(2) = 130.54, p < .001, Cramer V = 0.44. So smoking and COPD are
    not independent -- COPD is markedly more common among smokers. This is the statistical confirmation of one of
    medicine's best-known cause-effect relationships. For association between two categorical variables the
    chi-square test of independence is the standard method; Cramer V measures the strength. In epidemiology it is the
    basic tool for showing risk factor-disease relationships.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_blood_group.xlsx
  >> SCENARIO (narration):
    Imagine our infection ward wants to check whether the distribution of ABO
    blood groups among admitted patients matches the expected population
    proportions. We sampled 200 patients and recorded blood_group with four
    categories - 0, A, B, and AB. Rather than comparing two variables, we are
    testing whether one observed categorical distribution deviates from a
    hypothesized reference distribution. That is exactly what the Chi-Square
    Goodness-of-Fit test does, so we use it to see if our patient pool over- or
    under-represents any blood group relative to expectation.
  >> VARIABLE SELECTION:
    - Categorical variable: blood_group (0, A, B, AB)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 88.12   p < .001 ***   n = 200   k = 4 blood groups (expected: equal 1/k)
    H0 REJECTED

>> COMMENTARY (narration):
    We tested whether patients' blood groups (A, B, AB, O) are equally distributed with the chi-square goodness-of-fit
    test. The observed distribution deviates very significantly from the equal expectation (chi-square(3) = 88.12,
    p < .001): blood groups are not balanced, some far more common than expected. This is an expected result -- blood
    groups are naturally unequal in a population (e.g. A and O more common). The goodness-of-fit test is the right way
    to compare an observed category distribution against a theoretical expectation (equal or known population
    proportions); it is used in epidemiology for population-distribution comparisons.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_diagnosis_test.xlsx
  >> SCENARIO (narration):
    Suppose we are comparing a new rapid antigen test against PCR for confirming
    an infection in a small pilot of just 28 patients. The variable test marks
    which method was used - rapid test or PCR - and result records whether the
    outcome came back positive or negative. We want to know whether positivity
    depends on the testing method. With such a small sample and likely sparse
    cell counts, the chi-square approximation is unreliable, so Fisher's Exact
    Test is the appropriate choice for this 2x2 table.
  >> VARIABLE SELECTION:
    - Row variable: test (rapid test, PCR)
    - Column variable: result (positive, negative)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher Exact p = 0.276 ns   Odds Ratio = 0.39   Phi = 0.23 (Moderate)
    Diagnosis is NOT significantly associated with test result.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested the association between a diagnosis and a test result with Fisher's exact test -- because chi-square is
    unreliable for tables with few observations (small cells). This time the result is non-significant (p = 0.276):
    in this small sample no significant association between diagnosis and test result could be proven. A
    non-significant result is information too -- but we must keep in mind that power may be low in a small sample (a
    real association might go undetected). Fisher's exact test gives more accurate results than chi-square in
    small-sample/rare-outcome 2x2 tables; it is often needed in clinical pilot studies.

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_old_new_diagnosis.xlsx
  >> SCENARIO (narration):
    Here we evaluate whether a new diagnostic method changes how often infection
    is detected compared with the old culture-based method, using the same 90
    patients tested by both. For each patient, old_method and new_method each
    record absent or culture_positive. Because these are paired binary results
    on the same individuals, an ordinary chi-square would be wrong - we need to
    focus on the discordant pairs where the two methods disagree. The McNemar
    Test does exactly this, telling us whether the new method systematically
    shifts detection rates.
  >> VARIABLE SELECTION:
    - Pair measurement 1 (M1): old_method (absent, culture_positive)
    - Pair measurement 2 (M2): new_method (absent, culture_positive)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar (exact binomial) min(b,c) = 4   p = 0.077 ns   discordant: b = 12, c = 4
    H0 NOT REJECTED (no significant difference)

>> COMMENTARY (narration):
    We compared an old and a new diagnostic method on the same patients -- paired binary data call for McNemar's test.
    The result is borderline non-significant (p = 0.077): the new method caught what the old missed in 12 patients,
    while the old caught the new in 4 -- a trend favouring the new method but not crossing the significance threshold.
    The small number of discordant cells (b+c=16) limits the test's power. McNemar looks only at the cells that
    changed (discordant cells); it is the right way to compare two diagnostic methods on the same patients. A larger
    sample could clarify the new method's superiority.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_infection_stage_pathologist.xlsx
  >> SCENARIO (narration):
    Suppose two pathologists independently staged the infection severity of 70
    patients, each assigning one of four ordered stages from stage 1 to stage 4.
    We do not just want to know whether they agree by chance - we want a measure
    of how strongly their ratings coincide beyond what random agreement would
    produce. Cohen's Kappa is designed precisely for this inter-rater
    reliability question with two raters using the same categorical scale, so we
    use pathologist_a and pathologist_b to quantify their agreement.
  >> VARIABLE SELECTION:
    - Rater 1: pathologist_a (stage 1, stage 2, stage 3, stage 4)
    - Rater 2: pathologist_b (stage 1, stage 2, stage 3, stage 4)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's Kappa = 0.6705 (substantial/strong agreement)   p < .001 ***
    Two pathologists' infection staging

>> COMMENTARY (narration):
    We measured the agreement between two pathologists independently staging infection on the same samples with
    Cohen's Kappa. Kappa = 0.67 is in the "substantial/strong" range -- a chance-corrected measure of agreement, i.e.
    the real consistency after removing accidental agreement. In pathology this matters greatly: if staging is
    inconsistent between pathologists, diagnosis and treatment decisions vary by rater. Kappa = 0.67 shows the staging
    is largely reliable (with room for improvement). Reporting inter-rater reliability is the quality-assurance step
    of pathology/radiology studies.

====================================================================================

#26  Cochran-Mantel-Haenszel Test
    file: 26_cmh_treatment_septic_shock_age.xlsx
  >> SCENARIO (narration):
    Imagine a trial testing whether an antibiotic reduces progression to septic
    shock compared with placebo across 360 patients. The complication is that
    age strongly affects shock risk, so we stratify by age_group into three
    bands - under 55, 55 to 70, and over 70. We want the treatment-by-outcome
    association while controlling for age as a confounder. The Cochran-Mantel-
    Haenszel Test pools the 2x2 treatment-versus-septic_shock tables across the
    age strata, giving us a single adjusted association.
  >> VARIABLE SELECTION:
    - Row variable: treatment (placebo, antibiotic)
    - Column variable: septic_shock (absent, present)
    - Stratum / layer variable: age_group (<55, 55-70, >70)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi-square(1) = 12.62   p < .001 ***   Common OR (MH) = 3.80  (95% CI: 1.75 - 8.25)
    Breslow-Day chi-square(2) = 0.48, p = 0.79 (homogeneous OR)   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between treatment and septic shock while controlling for age groups (strata) with the
    CMH test. Even holding age constant the association is strong: common odds ratio 3.80 (p < .001) -- stripped of
    the age confounder, the odds of developing septic shock in one treatment group are 3.8 times the other's. The
    Breslow-Day test is non-significant (p = 0.79), meaning the association is consistent across all age groups
    (homogeneous OR). CMH is the classic way -- very valuable in clinical research -- to control for a third variable
    (age, centre) by stratification.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_sex_comorbidity_diabetes.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand how three patient characteristics interrelate
    in our cohort of 118 patients: sex (female or male), comorbidity burden
    coded as normal or high, and whether the patient has diabetes (absent or
    present). Rather than testing one pairwise association, we want to model the
    full pattern of dependencies among all three categorical variables at once,
    including possible higher-order interactions. Log-Linear Analysis is built
    for this, modeling the cell counts of the multi-way contingency table to
    reveal which associations are needed to explain the data.
  >> VARIABLE SELECTION:
    - Factor 1: sex (female, male)
    - Factor 2: comorbidity (normal, high)
    - Factor 3: diabetes (absent, present)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Best model AIC = 50.54   Pearson chi-square = 0.46 (model fits the data very well)
    Factors: sex x comorbidity x diabetes

>> COMMENTARY (narration):
    We analysed the contingency structure formed jointly by three categorical variables -- sex, comorbidity and
    diabetes -- with a log-linear model. The selected model fits very well (Pearson chi-square = 0.46, low), with AIC
    = 50.54 as the most parsimonious model. Log-linear analysis lets us untangle interactions among more than two
    categorical variables (which are dependent, which interaction is significant) -- it is like ANOVA for categorical
    data. In the clinic it is powerful for summarising multi-way relationships among comorbidity, sex and disease.

====================================================================================

#28  Crosstabulation
    file: 28_cross_age_symptom.xlsx
  >> SCENARIO (narration):
    Here we simply want to describe how presenting symptoms break down across
    age groups in 350 admitted patients. The variable age_group has three bands
    - 18-40, 41-65, and 66+ - and symptom records the chief complaint as abdomen
    pain, fever, or chest pain. Before any formal testing, clinicians often want
    a clear contingency table with counts and percentages to see, for example,
    whether fever dominates in younger patients. Crosstabulation gives us
    exactly that joint frequency layout of the two categorical variables.
  >> VARIABLE SELECTION:
    - Row variable: age_group (18-40, 41-65, 66+)
    - Column variable: symptom (abdomen pain, fever, chest pain)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(4) = 23.45   p < .001 ***   Cramer V = 0.18 (Weak-moderate)
    Age group is associated with symptom (weak).   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between age group and observed symptom with a crosstab and chi-square. The association
    is significant but weak (chi-square(4) = 23.45, p < .001, Cramer V = 0.18): age is related to the symptom
    distribution, but not strongly. Reading significance together with effect size matters: in a large sample even a
    weak association can be significant. Different age groups have slightly different symptom profiles, but age is not
    the sole determinant. The crosstab is the most basic and readable way to see the co-occurrence of two categorical
    variables.

====================================================================================

#29  Multiple Response (MR) Frequency
    file: 29_mr_frequency_comorbidity.xlsx
  >> SCENARIO (narration):
    Suppose each of our 250 patients can carry several comorbidities at once -
    the comorbidity field lists any combination such as diabetes, copd,
    hypertension, cardiovascular, anemia, asthma, rheumatic, or thyroid, with
    absent for those who have none. Because a single patient can select multiple
    conditions, a standard frequency table would miscount. The Multiple Response
    Frequency analysis treats comorbidity as a multiple-response set and tells
    us how often each individual condition appears across the cohort and what
    share of patients report it.
  >> VARIABLE SELECTION:
    - Multiple response set (items): comorbidity (e.g. diabetes, copd, hypertension, cardiovascular, anemia, asthma, rheumatic, thyroid; absent = none)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)
    Comorbidities (hypertension, diabetes, COPD, CKD...) selected as multi-choice

>> COMMENTARY (narration):
    Patients' comorbidities were recorded -- because a patient can have several coexisting conditions, this is a
    "multiple response" structure. All 250 patients carry at least one comorbidity, so the response rate is 100%.
    Multiple-response frequency analysis shows in how many patients and in what percentage each comorbidity occurs;
    the totals exceed 100% because each patient can have several conditions. Profiling the comorbidity burden of a
    patient population is very useful in clinical research for risk stratification and treatment planning.

====================================================================================

#30  Multiple Response Crosstab
    file: 30_mr_categorical_comorbidity_age.xlsx
  >> SCENARIO (narration):
    Now we want to see how the burden of multiple comorbidities differs by age
    in 220 patients. The comorbidity field is again a multiple-response set
    where a patient may list several of diabetes, copd, hypertension, sepsis and
    so on, while age_group splits patients into under 50, 50 to 70, and over 70.
    We want to cross each individual comorbidity against age band to spot, for
    instance, whether copd concentrates in the oldest group. The Multiple
    Response Crosstab pairs the multiple-response comorbidity set with the
    single categorical age_group variable.
  >> VARIABLE SELECTION:
    - Multiple response set (items): comorbidity (e.g. diabetes, copd, hypertension, sepsis; absent = none)
    - Crosstab (column) variable: age_group (<50, 50-70, >70)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: age group
    Comorbidities crossed by age group

>> COMMENTARY (narration):
    We cross-tabulated patients' comorbidities (multiple-choice) by age group. With data from 220 patients we can see
    how each comorbidity is distributed across age groups. The multiple-response crosstab answers "which age group
    carries which comorbidities more" -- e.g. hypertension/diabetes in the elderly group, fewer comorbidities in the
    young. This reveals age-related comorbidity patterns and guides age-specific clinical management.

====================================================================================

#31  MR x MR Crosstab
    file: 31_mr_mr_antibiotic_yanetki.xlsx
  >> SCENARIO (narration):
    Polypharmacy is common in our infectious-disease ward, and we want to know
    which drug classes tend to be co-prescribed with which adverse events. For
    180 patients we recorded two multiple-response variables: every drug class
    the patient was taking (antibiotics, which here spans wide-spectrum
    antibiotic, antihypertensive, Statin, antidiabetic and anticoagulant) and
    every adverse event they reported (side_effects: nausea, head pain, fatigue,
    skin reaction and fever). Because each patient can mark several drugs AND
    several side effects, an ordinary crosstab would miscount; a multiple-
    response by multiple-response (MR x MR) crosstab is the right tool. It lets
    us read, for each drug class, how often each adverse event appears, so we
    can flag pairings worth a closer pharmacovigilance look.
  >> VARIABLE SELECTION:
    - First multiple-response set (rows): antibiotics (wide_spectrum, antihypertensive, Statin, antidiabetic, anticoagulant)
    - Second multiple-response set (columns): side_effects (nausea, head pain, fatigue, skin, fever)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180
    Antibiotics x side effects (fatigue, fever, headache, nausea, skin) cross-tabulated

>> COMMENTARY (narration):
    This is the most advanced multiple-response analysis: we cross two separate multi-choice questions against each
    other -- the antibiotics given to the patient and the side effects observed. Among 180 patients we can see which
    antibiotics are associated with which side effects: e.g. one antibiotic goes more with nausea, another with skin
    reactions. This table reveals drug-side effect matches, providing powerful data for pharmacovigilance (drug
    safety monitoring) and treatment selection.

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q_symptom_time.xlsx
  >> SCENARIO (narration):
    We are following 150 patients through a 12-week antibiotic course and want
    to know whether a key clinical symptom truly resolves over time. At week 0,
    week 4 and week 12 we recorded for each patient whether the symptom was
    present (symptom_present: yes/no). Since the same binary outcome is measured
    on the same patients at three time points, the observations are dependent
    and the proportions are paired across time. Cochran's Q test is appropriate
    here because it tests whether the proportion of patients who are symptomatic
    changes across the three repeated binary measurements.
  >> VARIABLE SELECTION:
    - Repeated binary outcome: symptom_present (yes, no)
    - Within-subject time factor / repeated measures: time (week 0, week 4, week 12)
    - Subject identifier: patient_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000   (NO difference in symptom presence across time points)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the same patients' symptom presence (yes/no) across different time points with Cochran's Q test --
    the multi-occasion counterpart of repeated binary measures. The result is non-significant (Q = 0, p = 1.0): there
    is no difference in symptom prevalence across time points -- symptom presence stayed constant over time. A
    non-significant result is valuable too; it says the symptom does not change over time in this data (a chronic/
    stable course). Cochran's Q is the right way to test change in repeated binary (yes/no) measurements on the same
    patients; finding no expected change is a genuine result.

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5klinik.xlsx
  >> SCENARIO (narration):
    Before building any predictive model we want to understand how the core
    clinical and laboratory markers move together in our cohort of 180 patients.
    We have age, BMI, systolic_bp, lactate_mg_dL and ferritin_mg_dL, and we ask
    which of these are linearly associated, for example whether higher lactate
    tends to track with higher ferritin as a sign of more severe infection.
    Correlation analysis is the natural first step because all five variables
    are continuous and we simply want the pairwise strength and direction of
    their associations in a single correlation matrix.
  >> VARIABLE SELECTION:
    - Variables for correlation matrix: age
    - Variables for correlation matrix: BMI
    - Variables for correlation matrix: systolic_bp
    - Variables for correlation matrix: lactate_mg_dL
    - Variables for correlation matrix: ferritin_mg_dL

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables (age, BMI, systolic BP, lactate, ferritin)
    NO strong relationship (|r| >= 0.75) -- relationships are moderate/weak

>> COMMENTARY (narration):
    We computed all pairwise correlations among five key clinical variables. No pair exceeds the |r| >= 0.75
    threshold -- so the relationships are moderate or weak. This is typical in the clinic: age, BMI, blood pressure,
    lactate and ferritin may be interrelated but none alone determines another; clinical status is multi-factorial.
    The absence of strong relationships also means collinearity risk is low when building a regression. The
    correlation matrix is the first step before modelling to see the structure among variables and detect
    collinearity.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Analysis
    file: 34_bland_altman_bp_method.xlsx
  >> SCENARIO (narration):
    On the ward we are about to replace manual blood-pressure measurement with
    an automatic device, and before we trust it clinically we need to know
    whether the two methods agree. For 60 patients we measured systolic pressure
    twice, once manually (bp_manual) and once with the automatic monitor
    (bp_automatic), on the same patient. The clinical question is not whether
    the values correlate but whether they agree closely enough to be
    interchangeable, so Bland-Altman analysis is the right choice: it plots the
    difference between methods against their mean and gives us the bias and 95%
    limits of agreement.
  >> VARIABLE SELECTION:
    - Method 1 (Pair M1): bp_manual
    - Method 2 (Pair M2): bp_automatic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Manual BP mean = 128.60   |   Automatic BP mean = 127.33   |   Bias ~ +1.3 (methods agree largely)

>> COMMENTARY (narration):
    We compared two methods of measuring blood pressure -- manual (auscultation) and automatic device -- on the same
    patients with Bland-Altman. This method does not look at correlation (high correlation does not guarantee
    agreement); it measures the difference between the two measurements. The means are very close (128.6 vs 127.3),
    the bias about 1.3 mmHg -- the two methods give almost the same result systematically and are practically
    interchangeable. Bland-Altman is the standard way to assess the agreement of two measurement tools (manual/
    automatic, two devices, old/new method); it is indispensable in medical device validation.

====================================================================================

#35  Effect Size
    file: 35_effect_size_antidepresan.xlsx
  >> SCENARIO (narration):
    In a head-to-head comparison we randomized 120 patients to two antibiotic
    regimens and recorded how much their symptom score improved
    (symptom_score_change), with the regimen coded in antibiotic as SSRI versus
    SNRI. A p-value alone would only tell us whether a difference exists, but
    reviewers and clinicians want to know how large the benefit is in practice.
    Effect size analysis is appropriate because it quantifies the standardized
    magnitude of the difference in symptom_score_change between the two groups,
    giving a Cohen's d that is comparable across studies.
  >> VARIABLE SELECTION:
    - Group (2 levels): antibiotic (SSRI, SNRI)
    - Dependent variable: symptom_score_change

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.67  (moderate-large effect)
    Symptom-score change difference between two antibiotic groups

>> COMMENTARY (narration):
    We measured the symptom-score change difference between two treatment groups not just as "is it significant" but
    "how large is it" -- as an effect size. Cohen's d = -0.67 is a moderate-large effect. P-values inflate with
    sample size, but effect size is sample-independent and gives the clinical importance of the real difference. In
    medical research, "how much difference does this treatment make" is critical for clinical decisions and
    cost-effectiveness; a moderate-large effect of d = 0.67 says the treatment provides a notable benefit. Reporting
    effect size in publications and meta-analyses is standard.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_lab_clinical.xlsx
  >> SCENARIO (narration):
    We want to know whether a patient's overall laboratory profile on admission
    relates to their overall severity-score profile in 150 critically ill
    patients. One set of variables captures routine labs (WBC_K_uL, CRP_mg_L,
    creatinine_mg_dL, AST_U_L, Hb_g_dL) and the other set captures the clinical
    severity scores (APACHE_score, SOFA_score, SAPS_score, GCS). Instead of
    running many separate correlations, canonical correlation analysis lets us
    find the linear combinations of the lab set and the severity set that are
    maximally correlated, so we can describe the strongest shared dimension
    between biochemistry and clinical severity.
  >> VARIABLE SELECTION:
    - Set X (lab variables): WBC_K_uL
    - Set X (lab variables): CRP_mg_L
    - Set X (lab variables): creatinine_mg_dL
    - Set X (lab variables): AST_U_L
    - Set X (lab variables): Hb_g_dL
    - Set Y (severity scores): APACHE_score
    - Set Y (severity scores): SOFA_score
    - Set Y (severity scores): SAPS_score
    - Set Y (severity scores): GCS

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.970, chi-square(20) = 866.35, p < .001   |   CC2: r = 0.936, p < .001
    X set: WBC/CRP/creatinine/AST/Hb (lab)  |  Y set: APACHE/SOFA/SAPS/GCS (severity scores)

>> COMMENTARY (narration):
    We related two multivariate sets -- laboratory values (WBC, CRP, creatinine, AST, Hb) and ICU severity scores
    (APACHE, SOFA, SAPS, GCS) -- with canonical correlation. The first canonical axis is extremely strong (r = 0.97,
    p < .001), the second also significant (r = 0.94): the laboratory profile and disease severity are very tightly
    linked. So deteriorating lab values go together with high severity scores -- an expected but quantified clinical
    truth. CCA is the most direct answer to "how do two blocks of variables co-vary"; it is a powerful method for
    relating clinical biomarkers to disease severity.

====================================================================================

#37  Correspondence Analysis (CA)
    file: 37_ca_diagnosis_age.xlsx
  >> SCENARIO (narration):
    Across 550 admissions we want to see how the main admitting diagnoses map
    onto patient age bands. The diagnosis variable has four levels (MI,
    septic_shock, copd, diabetes) and age_group has three (18-40, 41-65, 66+),
    and we suspect, for instance, that septic_shock and COPD cluster in older
    patients while the pattern differs for the younger band. Correspondence
    analysis is suited to this because both variables are categorical; it turns
    the diagnosis-by-age contingency table into a low-dimensional map where rows
    and columns that lie close together are associated.
  >> VARIABLE SELECTION:
    - Row categorical variable: diagnosis (MI, septic_shock, copd, diabetes)
    - Column categorical variable: age_group (18-40, 41-65, 66+)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.046   Number of dimensions = 2
    Diagnosis x age group contingency table

>> COMMENTARY (narration):
    We projected the contingency table of diagnosis by age group into a two-dimensional map with correspondence
    analysis. Total inertia 0.046 indicates a relatively weak association -- diagnoses do not separate very sharply
    by age group. Diagnosis-age pairs that fall close on the map indicate that diagnosis is associated with that age
    group. The low inertia says there is no strong dependence between age and diagnosis (which is itself information).
    Correspondence analysis is a powerful way to visualise categorical relationships and also gives the association
    strength (inertia) numerically.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_SF36_18.xlsx
  >> SCENARIO (narration):
    We administered an 18-item quality-of-life questionnaire to 180 recovering
    patients, with six items each meant to tap a physical, a mental and a social
    dimension (physical_m1 to physical_m6, mental_m1 to mental_m6, social_m1 to
    social_m6). Before scoring we want to confirm empirically that the items
    group the way we intended and to spot any item that drifts to the wrong
    cluster. Variable clustering (VarClus) is appropriate because it groups the
    continuous items into clusters of mutually correlated variables, letting us
    check whether the data-driven clusters match the three theorized subscales.
  >> VARIABLE SELECTION:
    - Variables to cluster: physical_m1
    - Variables to cluster: physical_m2
    - Variables to cluster: physical_m3
    - Variables to cluster: physical_m4
    - Variables to cluster: physical_m5
    - Variables to cluster: physical_m6
    - Variables to cluster: mental_m1
    - Variables to cluster: mental_m2
    - Variables to cluster: mental_m3
    - Variables to cluster: mental_m4
    - Variables to cluster: mental_m5
    - Variables to cluster: mental_m6
    - Variables to cluster: social_m1
    - Variables to cluster: social_m2
    - Variables to cluster: social_m3
    - Variables to cluster: social_m4
    - Variables to cluster: social_m5
    - Variables to cluster: social_m6

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (SF-36 quality of life 18 items: physical / mental / social)
    A representative item was selected from each cluster

>> COMMENTARY (narration):
    We clustered the 18 items of the SF-36 quality-of-life scale by similarity into three groups. VarClus clusters
    items, not observations: it shows which items measure the same "construct". The three clusters formed as expected:
    physical, mental and social health items clustered separately. This confirms the scale really measures three
    sub-dimensions and lets us pick one representative per group to reduce the number of items. In examining the
    construct validity of quality-of-life scales and in item reduction, it is extremely practical.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_bp.xlsx
  >> SCENARIO (narration):
    We want to understand which patient characteristics jointly predict systolic
    blood pressure in a cohort of 200 patients. Our candidate predictors are
    age, BMI, daily fluid intake (fluid_g_day) and weekly mobilization time
    (mobilization_hour_week), and the outcome systolic_bp is continuous.
    Multiple linear regression is the right model here because it estimates the
    independent contribution of each predictor to systolic_bp while holding the
    others constant, so we can say, for example, how much pressure changes per
    unit of BMI after adjusting for age, fluids and mobilization.
  >> VARIABLE SELECTION:
    - Dependent variable: systolic_bp
    - Predictor: age
    - Predictor: BMI
    - Predictor: fluid_g_day
    - Predictor: mobilization_hour_week

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.683   Adjusted R2 = 0.676
    Systolic BP ~ age + BMI + fluid intake + mobilization

>> COMMENTARY (narration):
    We built a multiple linear regression predicting systolic blood pressure from four variables -- age, BMI, daily
    fluid intake, weekly mobilization. The model is strong: together they explain 68% of blood-pressure variance
    (R2 = 0.68). This is a good explanatory rate in the clinic -- blood pressure is multi-factorial and 68% captures a
    meaningful portion. Such a model shows which variable contributes most to blood pressure and helps evaluate the
    potential effect of lifestyle interventions (fluid, movement). Multiple regression is the basic tool for
    explaining a continuous clinical outcome with several drivers.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_comorbidity.xlsx
  >> SCENARIO (narration):
    In 300 patients we want to predict the probability of an organ attack
    (organ_attack, coded 0/1) from routinely available clinical data. The
    candidate predictors are age, ferritin, smoking status (smoking) and
    systolic_bp, and the outcome is binary. Logistic regression is appropriate
    because it models the log-odds of organ_attack as a function of these
    predictors, giving adjusted odds ratios that tell us how each marker, for
    example a rise in ferritin, changes the odds of the event while controlling
    for the rest.
  >> VARIABLE SELECTION:
    - Dependent variable (binary target, 0/1): organ_attack
    - Predictor: age
    - Predictor: ferritin
    - Predictor: smoking
    - Predictor: systolic_bp

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: Logistic Regression   Pseudo R2 = 0.095
    Organ involvement (1/0) ~ age + ferritin + smoking + systolic BP

>> COMMENTARY (narration):
    We built a logistic regression predicting whether a patient will have organ involvement from age, ferritin,
    smoking and blood pressure. Because the outcome is binary (present/absent), logistic regression is the right
    choice. Pseudo R2 = 0.095 shows the model partly explains the outcome -- the coefficients generally say the
    probability of organ involvement rises as age and ferritin increase. In medical research, logistic regression is
    the most common way to predict an outcome (complication, organ involvement, mortality) from risk factors; it
    gives each factor's effect on the odds in an interpretable way.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Poisson / Negative Binomial Regression
    file: 41_poisson_admission_count.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand what drives how often chronic infection
    patients keep coming back to the hospital over a year. For 180 patients we
    recorded their age, whether diabetes is present, and the
    annual_admission_count, which is the number of inpatient admissions each
    patient had during the study year. Because the outcome is a count of events
    rather than a continuous measurement, we model it with Poisson or Negative
    Binomial regression, asking whether older age and the presence of diabetes
    are associated with a higher admission rate. If the counts are over-
    dispersed, the Negative Binomial form lets us account for that extra
    variability rather than assuming the mean equals the variance.
  >> VARIABLE SELECTION:
    - Dependent variable (count): annual_admission_count
    - Predictor: age
    - Predictor: diabetes_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 590.52   Deviance = 228.88
    Annual admission count ~ age + diabetes presence   (Poisson)

>> COMMENTARY (narration):
    We built a Poisson regression predicting patients' annual hospital-admission count -- a count variable -- from age
    and diabetes presence. Count data (0,1,2,...) are not normally distributed, so we use Poisson rather than linear
    regression. The model typically shows admissions rising with age and diabetes. Deviance and AIC assess model fit.
    Count regression is the correct way to relate count outcomes in the clinic (admissions, attacks, complications)
    to risk factors; it is far more appropriate than a linear model.

====================================================================================

#42  Multinomial Logistic Regression
    file: 42_multinomial_infection_stage.xlsx
  >> SCENARIO (narration):
    Imagine we want to predict which clinical stage of infection a patient will
    present with, where infection_stage can be early, medium, or advanced. For
    250 patients we measured their age and a continuous biomarker level, and we
    want to know how these two predictors shift the odds of being in one stage
    versus another. Since the outcome is a nominal category with three unordered
    levels, ordinary logistic regression is not enough, so we use Multinomial
    Logistic Regression to estimate a separate set of effects for each stage
    relative to a reference category.
  >> VARIABLE SELECTION:
    - Dependent variable (3 nominal levels): infection_stage (early, medium, advanced)
    - Predictor: age
    - Predictor: biomarker

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 468.02   Reference class: 'advanced'
    Infection stage ~ age + biomarker

>> COMMENTARY (narration):
    We modelled patients' infection stage (several categories) from age and a biomarker with multinomial logistic
    regression. Because there are more than two unordered categories, this method fits: 'advanced' is the reference,
    and a separate equation is built for each other stage. Coefficients read as "how many times the odds of being
    early-stage rather than advanced change as the biomarker rises". This relates the stage distribution to
    continuous predictors -- a valuable tool in disease staging and prognosis prediction.

====================================================================================

#43  Ordinal Logistic Regression
    file: 43_ordinal_response_level.xlsx
  >> SCENARIO (narration):
    Here we are studying how patients respond to an antibiotic across a graded
    clinical scale, where response_level runs from nonresponsive, to partial, to
    good, to full recovery. For 200 patients we recorded the administered
    dose_mg and the patient's age, and we want to know whether a higher dose
    pushes patients toward better response categories. Because the outcome is
    ordered rather than just nominal, Ordinal Logistic Regression is the right
    tool, since it respects the natural ranking of the response levels and
    estimates a single set of effects under the proportional-odds assumption.
  >> VARIABLE SELECTION:
    - Dependent variable (ordinal): response_level (nonresponsive, partial, good, full)
    - Predictor: dose_mg
    - Predictor: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 344.47
    Response level (low<medium<high) ~ dose (mg) + age

>> COMMENTARY (narration):
    We built an ordinal logistic regression predicting treatment response level (low, medium, high -- ordered) from
    drug dose and age. These classes are ordered, and the ordinal model uses that order, making it more powerful and
    interpretable than multinomial. The coefficients give a patient's tendency to move to a higher response level as
    dose increases (dose-response relationship). For ordered categorical outcomes (response level, disease stage,
    Likert), ordinal logistic is the right method; it is widely used in pharmacology for dose-response modelling.

====================================================================================

#44  PLS Regression
    file: 44_pls_gene_progression.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict how far a patient's infection has progressed from
    their gene expression profile. For 200 patients we measured twelve genes,
    gene_01 through gene_12, along with a continuous progression_score that
    captures disease severity. These genes are highly correlated with one
    another, which makes ordinary regression unstable, so we use Partial Least
    Squares regression. PLS condenses the twelve correlated predictors into a
    few latent components that best explain the progression score, giving us a
    stable model even when the predictors overlap heavily.
  >> VARIABLE SELECTION:
    - Dependent variable: progression_score
    - Predictors: gene_01
    - Predictors: gene_02
    - Predictors: gene_03
    - Predictors: gene_04
    - Predictors: gene_05
    - Predictors: gene_06
    - Predictors: gene_07
    - Predictors: gene_08
    - Predictors: gene_09
    - Predictors: gene_10
    - Predictors: gene_11
    - Predictors: gene_12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2_train = 0.894   R2_CV (5-fold) = 0.880
    Disease progression score ~ 12 gene expression variables

>> COMMENTARY (narration):
    We predicted a disease progression score from 12 inter-correlated gene-expression variables with PLS (Partial
    Least Squares) regression. When gene expressions are highly correlated, ordinary regression becomes unstable; PLS
    solves this by reducing the variables to a few orthogonal components. The model is very strong both in training
    (R2 = 0.89) and cross-validation (R2_CV = 0.88) -- no overfitting, excellent generalisation. Genomic/biomarker
    data are typically high-dimensional and collinear; PLS keeps both predictive power and interpretability for such
    data.

====================================================================================

#45  Probit Regression
    file: 45_probit_antibiotic_dose_response.xlsx
  >> SCENARIO (narration):
    Let's say we are running a dose-response study for an antibiotic and want to
    know how the probability of a clinical response rises as the dose increases.
    For 180 patients we recorded the administered dose_mg and a binary response
    indicating whether the patient responded. Because the outcome is a yes-or-no
    event and we assume an underlying normally distributed latent tolerance, we
    fit a Probit Regression, which links the dose to the response probability
    through the normal cumulative distribution and lets us estimate the dose at
    which half the patients respond.
  >> VARIABLE SELECTION:
    - Dependent variable (binary): response
    - Predictor: dose_mg

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 118.36   Pseudo R2 (McFadden) = 0.535
    Response (1/0) ~ dose (mg)

>> COMMENTARY (narration):
    We modelled whether an antibiotic dose produces a response with probit regression. Probit models a binary outcome
    like logistic but assumes a normal-distribution curve (cumulative normal) -- the classic model for dose-response
    studies in toxicology and pharmacology. Pseudo R2 = 0.535 is very strong: dose largely explains the probability
    of response -- a classic dose-response relationship. The probit curve helps find the dose at which the response
    probability reaches 50% (ED50). Probit and logistic usually give similar results; in the dose-response context
    probit is traditional.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_hastane_expenditure.xlsx
  >> SCENARIO (narration):
    Imagine we want to model how much patients pay out of pocket during a
    hospital stay for an infection. For 200 patients we recorded their
    admission_day count, whether insurance is present, and the
    pocket_expenditure_TL. The catch is that many insured patients pay nothing,
    so the expenditure variable piles up at zero and is censored from below. A
    plain linear regression would be biased here, so we use Tobit Regression,
    which explicitly accounts for the lower censoring at zero while estimating
    how length of stay and insurance status affect out-of-pocket cost.
  >> VARIABLE SELECTION:
    - Dependent variable (left-censored at 0): pocket_expenditure_TL
    - Predictor: admission_day
    - Predictor: insurance_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 2868.75   Tobit (left-censored: out-of-pocket expenditure floored)
    Out-of-pocket health expenditure ~ admission days + insurance presence

>> COMMENTARY (narration):
    We modelled patients' out-of-pocket health expenditure from length of stay and insurance presence. The problem:
    many insured patients spend nothing out of pocket (zero) or expenditure piles up below a floor value -- this is
    censored data, and ordinary regression gives biased results. Tobit regression correctly models this left-censored
    structure (AIC = 2868.75). Typically expenditure rises with length of stay and falls with insurance. The
    statistically correct solution for working with censored/piled-up health-expenditure data is the Tobit model.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_dose_response.xlsx
  >> SCENARIO (narration):
    Suppose we want to estimate how an antibiotic dose affects a clinical
    response score, but we also want full uncertainty around our estimates
    rather than a single point value. For 100 patients we recorded the dose_mg,
    their age, and a continuous response_score. We use Bayesian Linear
    Regression, which treats the coefficients as distributions and returns
    credible intervals, letting us say directly how probable it is that
    increasing the dose improves the response while adjusting for age. This is
    especially useful with a modest sample, where prior information and
    posterior distributions give a more honest picture than classical p-values.
  >> VARIABLE SELECTION:
    - Dependent variable: response_score
    - Predictor: dose_mg
    - Predictor: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Conjugate Normal-Inverse-Gamma prior   posterior sigma2 = 27.25
    Response score ~ dose (mg) + age

>> COMMENTARY (narration):
    We used Bayesian linear regression to predict response score from dose and age. The Bayesian advantage: the
    result is not a single p-value but a probability distribution (posterior) for the parameters -- we get a credible
    interval for each coefficient and a "probability the effect is positive". The posterior error variance is sigma2 =
    27.25. The Bayesian method honestly expresses uncertainty and can incorporate prior clinical knowledge. It is
    valuable in pharmacological/clinical research with small samples or where information must be carried over from
    previous studies.

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear_albumin_age.xlsx
  >> SCENARIO (narration):
    Here we are looking at how serum albumin changes across the lifespan in
    infection patients, and we suspect the relationship is not a straight line.
    For 80 patients we recorded age_year and albumin_g_cm2, and a scatter
    suggests albumin rises early in life and then declines, a clearly curved
    pattern. Because a linear fit would miss this shape, we use Nonlinear
    Regression to fit a curve directly, estimating the parameters of a nonlinear
    function that describes how albumin tracks with age.
  >> VARIABLE SELECTION:
    - Dependent variable: albumin_g_cm2
    - Predictor: age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: y = K / (1 + exp(-r·(x-x0)))  (logistic growth)   R2 = 0.961
    Albumin ~ age

>> COMMENTARY (narration):
    We modelled how albumin (a bone/tissue marker) level changes with age. The relationship is not linear: with age
    it shows an S-curve (logistic) form saturating toward a ceiling (K). We modelled this with a logistic function
    and the fit is excellent (R2 = 0.96). The curve shows at which age the value changes fastest and when it
    approaches the ceiling. Nonlinear regression is the way to correctly model biological growth/maturation curves --
    bone mineralisation, growth, organ function; forcing a straight line would misrepresent the age-related change.

====================================================================================

#49  Ridge Regression
    file: 49_ridge_biomarker_progression.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict infection progression from a large panel of
    correlated biomarkers. For 200 patients we measured fifteen biomarkers,
    biomarker_p01 through biomarker_p15, and a continuous progression outcome.
    Many of these biomarkers move together, which causes severe
    multicollinearity and inflates ordinary regression coefficients. We use
    Ridge Regression, which adds an L2 penalty that shrinks the coefficients
    toward zero and stabilizes the estimates, giving a more reliable predictive
    model when the biomarkers are highly inter-correlated.
  >> VARIABLE SELECTION:
    - Dependent variable: progression
    - Predictors: biomarker_p01
    - Predictors: biomarker_p02
    - Predictors: biomarker_p03
    - Predictors: biomarker_p04
    - Predictors: biomarker_p05
    - Predictors: biomarker_p06
    - Predictors: biomarker_p07
    - Predictors: biomarker_p08
    - Predictors: biomarker_p09
    - Predictors: biomarker_p10
    - Predictors: biomarker_p11
    - Predictors: biomarker_p12
    - Predictors: biomarker_p13
    - Predictors: biomarker_p14
    - Predictors: biomarker_p15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Ridge (alpha = 1.0): R2 = 0.798   Adj. R2 = 0.782   n = 200
    Disease progression ~ 15 biomarkers (mutually correlated)

>> COMMENTARY (narration):
    Predicting disease progression from 15 highly correlated biomarkers, ordinary regression becomes unstable
    (collinearity inflates coefficients). Ridge regression solves this by shrinking all coefficients -- it does not
    zero them but limits their magnitude, keeping prediction stable (R2 = 0.80). Ridge is ideal when "I want to keep
    all biomarkers but there is collinearity". For correlated biomarker panels, it prevents overfitting and provides
    stable estimates.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso_30_pred_5_active.xlsx
  >> SCENARIO (narration):
    Imagine we have screened thirty candidate predictors and want to know which
    ones actually drive a patient's treatment_response, suspecting that only a
    handful truly matter. For 250 patients we recorded thirty variables, x01
    through x30, alongside the continuous treatment_response. We use Lasso
    Regression, whose L1 penalty drives the coefficients of irrelevant
    predictors exactly to zero, automatically selecting the small active subset
    that explains the response. This gives us a sparse, interpretable model that
    highlights the few biomarkers worth following up clinically.
  >> VARIABLE SELECTION:
    - Dependent variable: treatment_response
    - Predictors: x01
    - Predictors: x02
    - Predictors: x03
    - Predictors: x04
    - Predictors: x05
    - Predictors: x06
    - Predictors: x07
    - Predictors: x08
    - Predictors: x09
    - Predictors: x10
    - Predictors: x11
    - Predictors: x12
    - Predictors: x13
    - Predictors: x14
    - Predictors: x15
    - Predictors: x16
    - Predictors: x17
    - Predictors: x18
    - Predictors: x19
    - Predictors: x20
    - Predictors: x21
    - Predictors: x22
    - Predictors: x23
    - Predictors: x24
    - Predictors: x25
    - Predictors: x26
    - Predictors: x27
    - Predictors: x28
    - Predictors: x29
    - Predictors: x30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Lasso (alpha = 0.1): R2 = 0.723   Adj. R2 = 0.685   n = 250
    18 of 30 predictor coefficients shrunk to zero (automatic feature selection)

>> COMMENTARY (narration):
    In this data a few real predictors were mixed with many irrelevant variables. The strength of Lasso shows here:
    the penalty term drives irrelevant variables' coefficients exactly to zero -- automatic variable selection. 18 of
    30 predictors were eliminated, and the model stayed strong at R2 = 0.72. Lasso is extremely useful when you want
    to pick the true drivers from many candidate biomarkers -- in high-dimensional clinical/genomic data. While Ridge
    shrinks all variables, Lasso eliminates the irrelevant ones entirely.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_stress_procalcitonin_bp.xlsx
  >> SCENARIO (narration):
    Clinicians often suspect that psychological stress can drive up inflammatory
    markers, which in turn affect cardiovascular status. Here we ask whether
    perceived stress raises procalcitonin, and whether that elevated
    procalcitonin then pushes up systolic blood pressure in our cohort of 200
    patients. In other words, does procalcitonin act as the biological middle-
    man between stress and blood pressure? We use a mediation analysis because
    it lets us split the total effect of stress on systolic_bp into a direct
    path and an indirect path that runs through procalcitonin_ng_mL.
  >> VARIABLE SELECTION:
    - Predictor (X): stress_PSS
    - Mediator (M): procalcitonin_ng_mL
    - Outcome (Y): systolic_bp

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = -0.14   p = 0.607   95% CI (-0.69 , 0.36)   NOT SIGNIFICANT
    Path: stress (PSS) -> procalcitonin -> systolic BP

>> COMMENTARY (narration):
    We tested whether stress's effect on blood pressure passes through procalcitonin with mediation analysis. This
    time the indirect effect is non-significant (-0.14, p = 0.607, CI includes zero): stress's effect on blood
    pressure does not pass "through" procalcitonin. A non-significant mediation is information too -- the proposed
    mechanism (stress -> inflammation -> blood pressure) is not supported in this data; if stress affects blood
    pressure, it must go by another route. Mediation analysis can both confirm and refute a mechanism; a negative
    result prompts us to reconsider the assumed causal chain.

====================================================================================

#52  Path Analysis
    file: 52_path_health_behavior.xlsx
  >> SCENARIO (narration):
    We want to understand how an infection-prevention education program actually
    changes patient behavior. Our theory is a chain: giving patients information
    improves their attitude, a better attitude strengthens their intention to
    comply, and stronger intention finally translates into actual preventive
    behavior. With 220 patients measured on info, attitude, intention, and
    behavior, we want to test all of these links at once rather than one by one.
    We use path analysis because it estimates this whole system of directed
    relationships simultaneously and tells us how well the hypothesized model
    fits.
  >> VARIABLE SELECTION:
    - Exogenous variable: info
    - Endogenous variable: attitude
    - Endogenous variable: intention
    - Endogenous (final outcome): behavior
    - Path model: info -> attitude -> intention -> behavior

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.928   RMSEA = 0.200   (poor fit -- model should be revised)
    Model: info/attitude/intention -> behavior; attitude -> intention

>> COMMENTARY (narration):
    We tested a health-behavior model (info -> attitude -> intention -> behavior) with path analysis -- this is the
    classic "planned behavior" framework of health psychology. CFI = 0.93 is acceptable but RMSEA = 0.20 is high --
    the hypothesised model does not fully match the data; some paths may be missing or extra. This is informative
    too: it says the theoretical model should be revised. Path analysis is a powerful way to separate direct and
    indirect effects and test health-behavior theories; both good and poor fit guide us in improving the model.

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_patient_visit_procalcitonin.xlsx
  >> SCENARIO (narration):
    We are following patients on antibiotic therapy and measuring their
    procalcitonin repeatedly at four visits: baseline V0, then V3, V6 and V12.
    We want to know how procalcitonin changes over the course of treatment while
    properly accounting for the fact that the four measurements come from the
    same patient and are therefore correlated. We use a linear mixed model
    because it treats visit as a fixed effect for the time trend and adds a
    random effect for patient_id, so each patient gets their own baseline.
  >> VARIABLE SELECTION:
    - Dependent variable: procalcitonin
    - Fixed effect (Factor): visit (V0, V3, V6, V12)
    - Random effect / grouping: patient_id

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group variance (random intercept) = 34.76   ICC = 0.936
    Procalcitonin ~ visit  +  (1 | patient)

>> COMMENTARY (narration):
    Modelling the same patients' procalcitonin across visits, the within-patient repeated measures are not
    independent. The linear mixed model solves this by adding patient as a random effect. ICC = 0.94 is very high:
    94% of procalcitonin variance comes from between-patient differences; within-visit change is relatively small. So
    patients are consistent within themselves, and the main difference is between individuals. LMM is the correct,
    indispensable method for nested/repeated-measures data (patient>visit) -- it models each patient's own
    trajectory and resolves the independence violation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation_missing_patient.xlsx
  >> SCENARIO (narration):
    Real clinical datasets are messy, and ours is no exception: across 150
    patients some values for age, BMI, systolic_bp, procalcitonin and ferritin
    are missing. Simply dropping incomplete rows would shrink our sample and
    could bias the results if the data are not missing completely at random. We
    use multiple imputation to generate several plausible complete datasets,
    analyze each, and pool the results, which preserves the uncertainty
    introduced by the missing values instead of pretending we knew them exactly.
  >> VARIABLE SELECTION:
    - Variables to impute: age
    - Variables to impute: BMI
    - Variables to impute: systolic_bp
    - Variables to impute: procalcitonin
    - Variables to impute: ferritin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (number of imputations) = 5
    Missing procalcitonin values estimated from age/BMI/BP/ferritin

>> COMMENTARY (narration):
    Missing observations are inevitable in clinical data (tests not done, lost to follow-up); deleting them both
    loses data and biases results. Multiple imputation estimates the missing procalcitonin values from other
    variables, creating 5 separate completed datasets, runs the analysis in each and combines the results -- thereby
    also accounting for estimation uncertainty. Under the MAR (missing at random) assumption this is the modern
    missing-data standard. In longitudinal clinical studies it is the soundest way to cope with lost patient data.

====================================================================================

#55  Generalized Estimating Equations (GEE)
    file: 55_gee_procalcitonin_treatment.xlsx
  >> SCENARIO (narration):
    In this trial 300 longitudinal records track procalcitonin across repeated
    visits for patients on either a dense or a standard antibiotic regimen. We
    want to estimate the average treatment effect on procalcitonin over time
    across the whole population, not patient by patient. We use generalized
    estimating equations because they model the population-averaged response
    while accounting for the within-patient correlation across visits using
    patient_id as the cluster.
  >> VARIABLE SELECTION:
    - Dependent variable: procalcitonin
    - Predictor (Group): treatment (dense, standard)
    - Within-subject / time: visit
    - Cluster (subject id): patient_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coefficient = -0.907   p < .001 ***   QIC = 306.75
    Procalcitonin ~ visit  (repeated measures, clustered within patient)

>> COMMENTARY (narration):
    We took repeated (multi-visit) procalcitonin measurements from the same patients -- one patient's measurements
    are dependent. GEE estimates the relationship at the "population average" level in such clustered data,
    accounting for the correlation structure. The visit coefficient is significant and negative (-0.907, p < .001):
    procalcitonin falls markedly as visits progress -- showing the treatment brings the infection response under
    control. GEE is an alternative to LMM: rather than modelling random effects, it corrects the correlation among
    observations and gives the marginal (average) effect.

====================================================================================

#56  Generalized Linear Mixed Model (GLMM)
    file: 56_glmm_side_effect.xlsx
  >> SCENARIO (narration):
    We are monitoring how often patients report side effects after being
    switched to a new antibiotic, counting events over follow-up months. Because
    the outcome side_effect_count is a count and each patient contributes
    several repeated observations, an ordinary regression would not fit. We use
    a generalized linear mixed model with a Poisson-type response:
    new_antibiotic and time_month are fixed effects, and a random intercept for
    patient_id captures the fact that some patients are simply more prone to
    side effects than others.
  >> VARIABLE SELECTION:
    - Dependent variable (count): side_effect_count
    - Fixed effect: new_antibiotic
    - Fixed effect: time_month
    - Random effect / grouping: patient_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    new_antibiotic: coef = 0.541, p < .001 ***
    Poisson GLMM: side-effect count ~ month + new antibiotic  +  (1 | patient)

>> COMMENTARY (narration):
    We modelled side-effect counts (Poisson) with both month/new-antibiotic fixed effects and a patient random effect
    -- a GLMM: mixed model + count distribution. The new antibiotic effect is significant and positive (coef 0.54,
    p < .001): the new antibiotic is associated with more side effects. This is an important pharmacovigilance warning
    -- the new drug's efficacy advantage must be weighed against its side-effect burden. GLMM combines the strengths
    of LMM (which assumes normality) and GLM (which ignores clustering) for "repeated/nested count or binary data".
    It is the correct framework for monitoring repeated side-effect counts per patient.

====================================================================================

#57  Regularized Regression (Elastic Net)
    file: 57_elasticnet_risk_score.xlsx
  >> SCENARIO (narration):
    We have 40 candidate predictors (x01 through x40) and want to build a
    clinical risk_score model, but many of these predictors are correlated with
    each other and we worry about overfitting with so many variables. We use
    regularized regression with an elastic net because it combines ridge and
    lasso penalties, shrinking unhelpful coefficients toward zero while still
    handling groups of correlated predictors gracefully, giving us a stable and
    interpretable risk model.
  >> VARIABLE SELECTION:
    - Target: risk_score
    - Predictors: x01, x02, x03, x04, x05, x06, x07, x08, x09, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24, x25, x26, x27, x28, x29, x30, x31, x32, x33, x34, x35, x36, x37, x38, x39, x40

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.123   R2 (train) = 0.491
    Risk score ~ 40 predictors

>> COMMENTARY (narration):
    Elastic Net is a hybrid of Ridge and Lasso: it both shrinks coefficients (like Ridge) and zeros out the
    irrelevant ones (like Lasso). Here we predicted a clinical risk score from 40 predictors; Elastic Net keeps
    correlated variable groups together while eliminating the irrelevant ones. The model gives R2 = 0.49. When there
    are many collinear predictors -- biomarker panels, clinical scores, demographics -- Elastic Net is a balanced
    choice offering both selection and stability. It softens Lasso's "pick one at random from a correlated group"
    problem.

====================================================================================

#58  Robust Regression
    file: 58_robust_age_creatinine.xlsx
  >> SCENARIO (narration):
    We want to model how serum creatinine relates to patient age, but a few
    patients with severe renal complications show extreme creatinine values that
    could distort an ordinary least squares line. Rather than deleting these
    patients, we use robust regression because it down-weights outliers
    automatically, giving us an age-creatinine relationship that reflects the
    bulk of our 100 patients without being dragged around by a handful of
    extreme cases.
  >> VARIABLE SELECTION:
    - Dependent variable: creatinine_mg_dL
    - Predictor: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust intercept = 0.695   vs   OLS intercept = 0.993   (outliers distorted OLS)
    Creatinine ~ age

>> COMMENTARY (narration):
    This data had a few outliers -- e.g. patients with renal failure and very high creatinine. Ordinary regression
    (OLS) takes outliers fully into account and is pulled toward them: the intercept is 0.99 in OLS but 0.70 in the
    robust model; the marked gap means outliers distorted OLS. Robust regression (Huber-T) gives less weight to
    outlying observations to preserve the true trend. In clinical data, where outlier patient values are inevitable,
    robust regression gives a more reliable slope estimate than OLS.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_hastane_cost.xlsx
  >> SCENARIO (narration):
    Hospital costs for infectious-disease admissions are highly skewed: most
    patients are inexpensive, but a small number of complicated cases drive
    enormous bills. If we only model the mean we miss what is happening to those
    expensive patients. With 200 patients we use quantile regression to see how
    age and admission_day relate to total_cost_TL not just at the median but at
    the upper quantiles, revealing the drivers of high-cost stays.
  >> VARIABLE SELECTION:
    - Dependent variable: total_cost_TL
    - Predictor: age
    - Predictor: admission_day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    q = 0.10 slope = 0.167   |   q = 0.50 (median)   |   q = 0.90  (varying effect)
    Total cost ~ age + admission days

>> COMMENTARY (narration):
    Ordinary regression models only the mean; but age and length of stay may affect cost differently for low-cost and
    high-cost patients. Quantile regression models the lower (q=0.10, low cost), middle (q=0.50) and upper (q=0.90,
    high cost) quantiles separately. A steeper slope at the upper quantile means "in the most expensive patients,
    length of stay has the largest effect on cost". In health economics, when the extremes (most costly cases) matter
    rather than the average, quantile regression reveals what mean-based models miss.

====================================================================================

#60  ROC Curve Analysis
    file: 60_roc_biomarker_infection.xlsx
  >> SCENARIO (narration):
    We have a candidate biomarker and want to know how well it discriminates
    patients who truly have an infection from those who do not, across 250
    patients. We need to judge its diagnostic accuracy and find a useful cut-
    off. We use ROC curve analysis because it plots sensitivity against
    1-specificity over every possible threshold of biomarker_score and
    summarizes overall performance with the area under the curve, using
    infection_present as the true disease status.
  >> VARIABLE SELECTION:
    - Test variable / predictor: biomarker_score
    - State variable (positive = infection): infection_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.969 (excellent discriminating power)   Youden optimum threshold: 0.47 -> Sensitivity 0.92
    n = 250 (positive 93, negative 157)   infection classifier

>> COMMENTARY (narration):
    We measured how well a biomarker score discriminates infection with the ROC curve. AUC = 0.97 is "excellent": the
    biomarker separates patients with and without infection almost perfectly. The Youden index gives the best cut-off
    (threshold), optimising sensitivity and specificity; here sensitivity is 0.92 at a 0.47 threshold. ROC/AUC is the
    gold standard for reporting a diagnostic test's/biomarker's discriminating power; it is the basis for evaluating
    the clinical value of medical diagnostic tools (biomarkers, screening tests).

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss_disease_prediction.xlsx
  >> SCENARIO (narration):
    Suppose we have built a screening model that flags patients at risk of a
    serious infection, and we want to know how well its predictions actually
    separate the sick from the healthy. In this cohort of 200 patients we
    recorded two clinical risk factors, the model's continuous prediction_score,
    and the ground-truth disease_present indicator confirmed at follow-up. The
    True Skill Statistic is ideal here because it summarizes sensitivity plus
    specificity minus one into a single threshold-independent score, so it isn't
    fooled when most patients are disease-free, which is exactly the prevalence
    imbalance we see in infection screening.
  >> VARIABLE SELECTION:
    - Predicted score / probability: prediction_score
    - Observed binary outcome: disease_present (0 = absent, 1 = present)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.389 (acceptable)   Sensitivity 0.44   Specificity 0.95   n = 200
    Disease prediction

>> COMMENTARY (narration):
    TSS (True Skill Statistic) measures how much better than chance a binary classification model is: sensitivity +
    specificity - 1. The model predicting disease presence has TSS = 0.39 -- "acceptable": specificity is very high
    (0.95, correctly clears the healthy) but sensitivity is low (0.44, misses many true cases). This balance is
    clinically important: the model is good at clearing non-cases but misses some real cases. TSS, being insensitive
    to prevalence, is a reliable performance measure for imbalanced disease data; it summarises the true skill of
    screening tests.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity_diabetes_class.xlsx
  >> SCENARIO (narration):
    Imagine we deployed an algorithm to classify whether incoming patients have
    diabetes-related complications, and clinicians need to know not just overall
    accuracy but where the model makes its mistakes. For 200 patients we have
    the actual_diagnosis label, the model's prediction_diagnosis label, and the
    underlying prediction_score. Confusion Matrix Metrics are the right tool
    because they break performance down into true positives, false positives,
    sensitivity, specificity, precision and F1, which lets us see whether the
    model is missing real cases or raising too many false alarms before we trust
    it at the bedside.
  >> VARIABLE SELECTION:
    - Actual / true class: actual_diagnosis
    - Predicted class: prediction_diagnosis
    - Predicted probability (optional, for threshold metrics): prediction_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.94   F1 = 0.923
    Diagnosis-class prediction -- actual vs predicted

>> COMMENTARY (narration):
    We broke down a diagnostic classifier's predictions with a confusion matrix. With accuracy 94% and F1 = 0.92 the
    model is very successful -- both the overall correct diagnosis rate is high and the precision/recall balance (F1)
    is good. The confusion matrix reveals what the "accuracy" number hides -- where each diagnostic class is
    misclassified, which diagnoses are confused. The F1 score is more informative than accuracy especially with
    imbalanced classes. It is the indispensable tool for evaluating the real performance of medical diagnostic/
    classification models (AI-assisted diagnosis).

====================================================================================

#63  Random Forest
    file: 63_rf_infection_type.xlsx
  >> SCENARIO (narration):
    Suppose an infectious-disease unit wants to predict which of four infection
    sites a patient is most likely to have from routine admission data, so
    empiric therapy can be started faster. In these 350 patients we have age,
    sex (sex_E), procalcitonin, creatinine and hemoglobin (Hb_g_dL) as
    predictors, and the confirmed infection_type as the target with levels
    colon, prostate, breast and lung. Random Forest fits well here because it
    handles nonlinear relationships and interactions among biomarkers without
    strong distributional assumptions, and it gives us a variable-importance
    ranking so we can see which labs drive the classification.
  >> VARIABLE SELECTION:
    - Target (4 classes): infection_type (colon, prostate, breast, lung)
    - Predictors: age
    - Predictors: sex_E
    - Predictors: procalcitonin
    - Predictors: creatinine
    - Predictors: Hb_g_dL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.90   (infection-type classification)
    Predicted from 5 clinical variables (age, sex, procalcitonin, creatinine, Hb)

>> COMMENTARY (narration):
    We predicted infection type from five clinical variables (age, sex, procalcitonin, creatinine, Hb) with Random
    Forest -- the vote of hundreds of decision trees. Accuracy is 90%: clinical variables largely classify the
    infection type correctly. One of Random Forest's most valuable outputs is the "variable importance ranking",
    telling which biomarker matters most for the diagnostic distinction. It is a powerful, easy-to-build machine-
    learning method for nonlinear, interacting clinical relationships; it forms the basis of diagnostic decision
    support systems.

====================================================================================

#64  Support Vector Machine (SVM)
    file: 64_svm_sepsis.xlsx
  >> SCENARIO (narration):
    Suppose we want an early-warning classifier that decides whether a patient
    is septic from a handful of bedside labs and vitals. For 250 patients we
    measured CRP (CRP_mg_L), white blood cell count (WBC_K_uL), peak fever in
    Celsius (fever_C) and lactate (lactate_mmol), with sepsis recorded as the
    binary outcome. A Support Vector Machine is a strong choice here because it
    finds the optimal margin separating septic from non-septic patients and,
    with a kernel, can capture the nonlinear boundary that often exists between
    inflammatory markers and clinical sepsis.
  >> VARIABLE SELECTION:
    - Target (binary): sepsis (0 = no, 1 = yes)
    - Predictors: CRP_mg_L
    - Predictors: WBC_K_uL
    - Predictors: fever_C
    - Predictors: lactate_mmol

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (sepsis present/absent)
    Predicted from 4 variables (CRP, WBC, fever, lactate)

>> COMMENTARY (narration):
    We found the best boundary separating patients with and without sepsis with SVM. SVM seeks the maximum-margin
    separating surface between two classes and, via the kernel trick, can model nonlinear boundaries too. Accuracy is
    100% -- CRP, WBC, fever and lactate separate sepsis almost perfectly (such high accuracy should be kept in mind
    regarding overfitting). SVM is especially powerful in small-to-medium, high-dimensional clinical classification
    problems. For early detection of conditions requiring urgent intervention like sepsis, it is a strong alternative
    to Random Forest.

====================================================================================

#65  Gradient Boosting
    file: 65_gb_repeat_admission.xlsx
  >> SCENARIO (narration):
    Imagine we want to stratify discharged infection patients by their risk of
    being readmitted, so the high-risk ones can get closer follow-up. In this
    dataset of 280 patients we have age, comorbidity_count, days since the last
    admission (last_admission_day) and the number of antibiotics received
    (antibiotic_count), and the outcome repeat_admission_risk is graded as low,
    medium or high. Gradient Boosting suits this problem because it builds an
    ensemble of trees sequentially, correcting previous errors, which typically
    gives strong predictive accuracy on tabular clinical data with mixed risk
    signals.
  >> VARIABLE SELECTION:
    - Target (3 classes): repeat_admission_risk (low, medium, high)
    - Predictors: age
    - Predictors: comorbidity_count
    - Predictors: last_admission_day
    - Predictors: antibiotic_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.893   (repeat-admission risk classification)
    age + comorbidity count + last admission day + antibiotic count

>> COMMENTARY (narration):
    We predicted repeat-admission (readmission) risk from clinical variables with Gradient Boosting. Boosting adds
    trees sequentially, each new tree correcting the previous ones' errors; accuracy is high at 89.3%. This is
    critical for health systems: high-readmission-risk patients can be identified in advance and discharge planning
    and follow-up intensity adjusted. While Random Forest builds trees in parallel, Boosting builds them sequentially
    (chasing the error); it usually gives higher accuracy but needs careful tuning. It is a valuable tool for hospital
    resource management.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_patient_4kume.xlsx
  >> SCENARIO (narration):
    Suppose we want to discover natural patient subgroups among infection
    admissions without using any predefined diagnosis, purely from their
    physiology. For 240 patients we have age, BMI, systolic_bp and lactate, and
    we suspect there may be roughly four phenotypes ranging from stable to
    critically ill. K-Means Clustering is appropriate because it partitions
    patients into a chosen number of compact clusters based on overall
    similarity in these continuous measures, giving us data-driven severity
    profiles we can then characterize clinically.
  >> VARIABLE SELECTION:
    - Clustering variables: age
    - Clustering variables: BMI
    - Clustering variables: systolic_bp
    - Clustering variables: lactate
    - Number of clusters (k): 4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.442 (moderate cluster structure)
    Patient typology from 4 variables (age, BMI, systolic BP, lactate)

>> COMMENTARY (narration):
    We split patients into four natural groups (phenotypes) by their clinical features with K-Means. Silhouette =
    0.44 is moderate -- the clusters separate but not sharply, which is expected in clinical data. The four groups
    likely represent distinct patient phenotypes: e.g. elderly-hypertensive, young-stable, high-lactate critical, and
    intermediate types. K-Means finds natural groups in unlabeled data; it is practical for patient phenotyping, risk
    stratification and forming personalised treatment groups. The silhouette score shows the number of clusters is
    reasonable.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5tip_10gen.xlsx
  >> SCENARIO (narration):
    Imagine we have molecular profiles from 40 patients spanning five suspected
    infection subtypes and we want to see whether their gene-expression patterns
    group together in a way that mirrors those subtypes. We have ten gene-
    expression features, gene_01 through gene_10, and a tentative lower_type
    label with levels type-A through type-E that we can use afterward to
    interpret the branches. Hierarchical Clustering is the right approach
    because it builds a dendrogram showing how patients merge step by step,
    letting us inspect the nested structure and decide the natural number of
    groups rather than fixing it in advance.
  >> VARIABLE SELECTION:
    - Clustering variables: gene_01
    - Clustering variables: gene_02
    - Clustering variables: gene_03
    - Clustering variables: gene_04
    - Clustering variables: gene_05
    - Clustering variables: gene_06
    - Clustering variables: gene_07
    - Clustering variables: gene_08
    - Clustering variables: gene_09
    - Clustering variables: gene_10
    - Label for interpreting clusters (not used in fitting): lower_type (type-A...type-E)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 5   Silhouette = 0.147 (weak/overlapping structure)
    5 types from 10 gene expressions

>> COMMENTARY (narration):
    We grouped patients by gene-expression profile with hierarchical clustering -- the result is a dendrogram (tree).
    Splitting into five clusters gives a silhouette of only 0.15: the clusters overlap, there is no sharp separation.
    This is common in genomic data -- gene-expression profiles often change along a continuous gradient rather than
    in sharp subtypes. A low silhouette means "natural grouping in the data is weak", which is informative (perhaps
    fewer than 5 molecular subtypes exist). Hierarchical clustering is ideal for discovering the nested structure of
    groups and the right number of subtypes.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan_case_location.xlsx
  >> SCENARIO (narration):
    Suppose we are mapping where confirmed infection cases occurred across a
    region and we want to find dense outbreak clusters while flagging isolated,
    scattered cases as background noise. We have 140 cases, each with a latitude
    (lat) and longitude (lon). DBSCAN is well suited here because, unlike
    k-means, it doesn't require us to guess the number of clusters and it
    naturally labels low-density points as outliers, which is exactly what we
    want when distinguishing genuine spatial hotspots from sporadic infections.
  >> VARIABLE SELECTION:
    - Coordinates (X): lon
    - Coordinates (Y): lat

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.911   (case locations)
    Cases clustered by geographic density

>> COMMENTARY (narration):
    We clustered infection cases' geographic locations by density with DBSCAN. Unlike K-Means: we do not give the
    number of clusters in advance, and it flags low-density isolated cases as "noise" (outliers). Four dense case
    groups were found (silhouette 0.91, very strong). This shows cases cluster in particular areas -- a critical
    finding in outbreak epidemiology (a possible common source, transmission focus). DBSCAN is more appropriate than
    K-Means for spatial case data with irregularly shaped clusters and outliers; it is used to detect outbreak
    hot-spots.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca_clinical_6ozellik.xlsx
  >> SCENARIO (narration):
    Imagine we have six correlated clinical measurements on each infection
    patient and we want to compress them into a few independent axes that
    summarize overall illness severity for easier visualization and downstream
    modeling. For 180 patients we recorded age, BMI, systolic_bp, lactate,
    ferritin and procalcitonin. Principal Component Analysis fits because these
    variables are intercorrelated, and PCA rotates them into orthogonal
    components that capture most of the variance, so we can often describe a
    patient with two or three composite scores instead of six raw values.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: BMI
    - Variables: systolic_bp
    - Variables: lactate
    - Variables: ferritin
    - Variables: procalcitonin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 = 91.9%   (the first component alone explains nearly all the variance)
    6 clinical features (age, BMI, systolic BP, lactate, ferritin, procalcitonin)

>> COMMENTARY (narration):
    We reduced six clinical features to a few summary axes with PCA. The first component alone explains 91.9% of the
    variance -- strikingly high. The meaning: these clinical features are very highly correlated, all measuring a
    single "general disease severity" dimension. So when one feature deteriorates in a patient, the others do too.
    PCA is the basic tool for summarising multivariate clinical data, reducing collinearity and revealing hidden
    dimensions (like general severity). One component being this dominant suggests a composite severity index could
    be built.

====================================================================================

#70  t-SNE
    file: 70_tsne_infection_5tip.xlsx
  >> SCENARIO (narration):
    Suppose we have high-dimensional gene-expression data from 125 patients
    across five infection subtypes and we want a two-dimensional map that
    reveals whether those subtypes form distinct groups. We have twelve gene
    features, gene_01 through gene_12, plus a lower_type label with levels A
    through E to color the points. t-SNE is the right tool because it is a
    nonlinear embedding that preserves local neighborhoods, so patients with
    similar expression profiles sit close together, making any subtype structure
    visually obvious even when it is hidden in the raw twelve-dimensional space.
  >> VARIABLE SELECTION:
    - Features: gene_01
    - Features: gene_02
    - Features: gene_03
    - Features: gene_04
    - Features: gene_05
    - Features: gene_06
    - Features: gene_07
    - Features: gene_08
    - Features: gene_09
    - Features: gene_10
    - Features: gene_11
    - Features: gene_12
    - Color / label (not used in fitting): lower_type (A, B, C, D, E)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence (fit quality) = 0.558 -> good
    12 gene expressions (5 infection types) reduced to 2D

>> COMMENTARY (narration):
    t-SNE is a nonlinear method that projects high-dimensional gene-expression data (12 genes) into a two-dimensional
    plot. Unlike PCA, it focuses on preserving local neighbourhoods -- similar patients are placed close on the map
    -- ideal for seeing hidden molecular-subtype structure. The KL divergence is 0.56, so the reduction quality is
    good; the five infection types form visible clusters on the map. Important caveat: t-SNE axes and between-cluster
    distances are not interpretable, only the clustering pattern is. It is powerful for visually exploring genomic/
    transcriptomic patient subtypes.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds_patient_distance.xlsx
  >> SCENARIO (narration):
    Suppose we want to map how septic patients relate to one another based on
    their overall clinical profile rather than any single marker. For 51
    patients we recorded age, BMI, systolic blood pressure, ferritin and
    procalcitonin, and we'd like to see which patients cluster together and how
    far apart their profiles are. Multidimensional Scaling is ideal here because
    it takes the pairwise distances between these multivariate clinical profiles
    and projects them into a low-dimensional map we can actually look at,
    preserving the relative dissimilarities so that patients with similar
    inflammatory and hemodynamic signatures land close together.
  >> VARIABLE SELECTION:
    - Coordinates / Input variables: age
    - Coordinates / Input variables: BMI
    - Coordinates / Input variables: systolic_bp
    - Coordinates / Input variables: ferritin
    - Coordinates / Input variables: procalcitonin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.091 (acceptable fit)
    Patient similarity map (age, BMI, BP, ferritin, procalcitonin)

>> COMMENTARY (narration):
    MDS maps the clinical similarity among patients into two dimensions -- patients with similar profiles placed near,
    dissimilar ones far. Stress = 0.091 is an "acceptable" fit; the patient-similarity structure largely fits two
    dimensions (low stress = reliable map). Patients close on the map have similar clinical profiles. MDS is the
    classic way to visualise multivariate similarity/distance data; the stress value honestly tells how much we can
    trust the map. It is used to see patient grouping and clinical-profile similarity.

====================================================================================

#72  UMAP
    file: 72_umap_omics_5tip.xlsx
  >> SCENARIO (narration):
    Imagine we've profiled 200 infection patients across 25 omics features
    (omics_01 through omics_25) and we suspect they fall into five biological
    subtypes labelled S1 to S5. The high-dimensional omics space is impossible
    to inspect directly, so we want a faithful two-dimensional embedding that
    preserves the local neighbourhood structure. UMAP is the right tool because
    it captures the manifold geometry of these omics signatures and reveals
    whether the five subtypes form distinct, well-separated clusters; we feed in
    the 25 omics features and color the embedding by the known subtype to
    validate the separation.
  >> VARIABLE SELECTION:
    - Features / Input variables: omics_01 ... omics_25 (all 25 omics columns)
    - Color / Grouping label: lower_type (S1, S2, S3, S4, S5)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    UMAP dimension reduction (25 omics variables -> 2D, 5 types)
    n_neighbors and min_dist balance local/global structure

>> COMMENTARY (narration):
    UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure
    better and is faster. We projected patient profiles of 25 omics variables (gene/protein/metabolite) into a
    two-dimensional map; similar molecular profiles cluster. The n_neighbors parameter tunes the local-global
    balance, and min_dist the tightness of clusters. UMAP is increasingly preferred for discovering hidden patient
    subtypes in high-dimensional omics data -- an important visualisation tool in precision-medicine studies.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_BDI_21.xlsx
  >> SCENARIO (narration):
    Suppose we administered a 21-item depression inventory (the BDI) to 200
    infection patients, with items item_01 through item_21 all scored on the
    same scale, and we want to confirm that these items reliably measure a
    single underlying construct before we sum them into a total score.
    Cronbach's Alpha is the appropriate reliability coefficient here because it
    quantifies the internal consistency among the 21 items, telling us how well
    they hang together and whether dropping any single item would meaningfully
    improve the scale.
  >> VARIABLE SELECTION:
    - Items: item_01
    - Items: item_02
    - Items: item_03
    - Items: item_04
    - Items: item_05
    - Items: item_06
    - Items: item_07
    - Items: item_08
    - Items: item_09
    - Items: item_10
    - Items: item_11
    - Items: item_12
    - Items: item_13
    - Items: item_14
    - Items: item_15
    - Items: item_16
    - Items: item_17
    - Items: item_18
    - Items: item_19
    - Items: item_20
    - Items: item_21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.974 (excellent internal consistency)   21 items
    Beck Depression Inventory (BDI)-like scale

>> COMMENTARY (narration):
    We tested the internal consistency of a 21-item depression/symptom scale (BDI-like) with Cronbach's alpha. Alpha
    = 0.97 is excellent: the 21 items consistently measure the same underlying concept (depression severity), and
    patients answer coherently across items. Showing a scale's reliability before analysing it is essential -- low
    alpha signals the items do not measure the same thing. Above 0.70 is acceptable, above 0.90 excellent. It is the
    first quality-control step of patient-reported outcome scales (PROMs) in the clinic.

====================================================================================

#74  Likert Analysis
    file: 74_likert_survival_quality_3boyut.xlsx
  >> SCENARIO (narration):
    Suppose we surveyed 220 patients on their survival-related quality of life
    using a Likert questionnaire organized into three dimensions: physical
    (physical_1 to physical_5), mental (mental_1 to mental_5) and social
    (social_1 to social_5). We want to summarize the response distribution for
    each item, see how respondents lean across the agreement categories, and
    compare the three dimensions. Likert Analysis is the right approach because
    it treats these ordinal items appropriately, producing per-item frequency
    profiles and diverging summaries that make the patients' perceived quality
    of life easy to interpret.
  >> VARIABLE SELECTION:
    - Likert items: physical_1
    - Likert items: physical_2
    - Likert items: physical_3
    - Likert items: physical_4
    - Likert items: physical_5
    - Likert items: mental_1
    - Likert items: mental_2
    - Likert items: mental_3
    - Likert items: mental_4
    - Likert items: mental_5
    - Likert items: social_1
    - Likert items: social_2
    - Likert items: social_3
    - Likert items: social_4
    - Likert items: social_5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.782 (acceptable)   15 items (3 dimensions: physical/mental/social)

>> COMMENTARY (narration):
    We examined the overall reliability and item statistics of a 15-item, three-dimensional (physical, mental, social)
    quality-of-life Likert scale. Cronbach's alpha = 0.78 is "acceptable" -- the scale is generally consistent, but
    because three different dimensions are combined, a single alpha can come out a bit low (sub-dimensions should be
    analysed separately). Likert analysis gives item means, distributions and item-total correlations to show whether
    the scale is sound. It is used for the basic quality control of quality-of-life/symptom scales in the clinic.

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18_item_3faktor.xlsx
  >> SCENARIO (narration):
    Imagine we developed an 18-item patient questionnaire (q01 through q18) and
    we don't yet know how many latent factors underlie the responses from our
    250 infection patients. Before naming or scoring any subscales, we want the
    data to tell us the factor structure. Exploratory Factor Analysis is
    appropriate because it uncovers the number of underlying latent dimensions
    and shows which items load on which factor, helping us discover whether the
    18 items resolve into the three constructs we hypothesized.
  >> VARIABLE SELECTION:
    - Items: q01
    - Items: q02
    - Items: q03
    - Items: q04
    - Items: q05
    - Items: q06
    - Items: q07
    - Items: q08
    - Items: q09
    - Items: q10
    - Items: q11
    - Items: q12
    - Items: q13
    - Items: q14
    - Items: q15
    - Items: q16
    - Items: q17
    - Items: q18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO (overall) = 0.909 (Excellent)   -> factor analysis very appropriate
    18 items, expected 3 factors

>> COMMENTARY (narration):
    We explored how many hidden dimensions (factors) underlie 18 scale items with Exploratory Factor Analysis. KMO =
    0.91 is "excellent": the data are very suitable for factor analysis, with high shared variance among items. The
    items group into three factors as expected. EFA reduces many items to a few meaningful dimensions, simplifying
    the scale and revealing the underlying structures. It is a core tool of clinical scale development and
    psychometric validation; KMO and Bartlett's test are reported as the pre-check of suitability.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3radyolog_tumor.xlsx
  >> SCENARIO (narration):
    Suppose three radiologists independently measured tumor size in millimeters
    for the same 50 patients, giving us radiologist_1_mm, radiologist_2_mm and
    radiologist_3_mm. Before we trust these measurements in a study, we need to
    know how consistently the raters agree on the same continuous quantity. The
    Intraclass Correlation Coefficient is exactly the right statistic because it
    assesses absolute agreement among multiple raters measuring the same
    patients, telling us whether the radiologists' tumor measurements are
    reliable enough to be used interchangeably.
  >> VARIABLE SELECTION:
    - Raters / Repeated measures: radiologist_1_mm
    - Raters / Repeated measures: radiologist_2_mm
    - Raters / Repeated measures: radiologist_3_mm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.951   (Excellent)   3 radiologists   95% CI ~ (0.92 , 0.97)

>> COMMENTARY (narration):
    We assessed how consistent three radiologists are in measuring tumour size (mm) on the same images with ICC. ICC
    = 0.95 is "excellent": inter-radiologist agreement is very high -- whichever radiologist measures, the result is
    similar. This is critical for clinical decision-making: if tumour-size measurement varied from radiologist to
    radiologist, the treatment decision (surgical margin, staging) would be unreliable. ICC is the standard for
    measuring inter-rater reliability on continuous measurements (size, score) (Cohen's Kappa is for categorical
    data, ICC for continuous). It is essential for radiological/laboratory measurement reliability.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_DASS21_12.xlsx
  >> SCENARIO (narration):
    Suppose we administered the 12-item DASS-21 short form to 280 infection
    patients, and from theory we expect three correlated latent factors: a
    symptom factor (symptom_1 to symptom_4), a fatigue factor (fatigue_1 to
    fatigue_4) and a stress factor (stress_1 to stress_4). Unlike an exploratory
    analysis, here we already have a hypothesized structure we want to test.
    Confirmatory Factor Analysis is the correct method because it lets us
    specify this three-factor model in advance and evaluate model fit,
    confirming whether each set of items loads cleanly on its intended factor.
  >> VARIABLE SELECTION:
    - Factor 1 (Symptom) indicators: symptom_1, symptom_2, symptom_3, symptom_4
    - Factor 2 (Fatigue) indicators: fatigue_1, fatigue_2, fatigue_3, fatigue_4
    - Factor 3 (Stress) indicators: stress_1, stress_2, stress_3, stress_4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 1.00   RMSEA = 0.000   (excellent fit)
    3 factors: symptom + fatigue + stress (DASS-like)

>> COMMENTARY (narration):
    While EFA explores hidden structure, CFA tests a pre-specified theory: here we tested whether three distinct
    dimensions -- "symptom", "fatigue" and "stress" (similar to the DASS scale) -- fit the data. The fit is excellent
    (CFI = 1.00, RMSEA = 0.000): the items really load onto these three dimensions, the hypothesised scale structure
    is fully confirmed. CFA is part of structural equation modelling (SEM) and is the strongest way to prove scale
    validity. In the clinic it is the standard method for demonstrating the validity of multidimensional
    patient-reported scales (depression, anxiety, stress).

====================================================================================

#78  Survey Means
    file: 78_survey_means_bp.xlsx
  >> SCENARIO (narration):
    Suppose our 350 patients were drawn from a survey design and we want a
    population-level estimate of mean systolic blood pressure, broken down by
    age group (18-40, 41-65, and 66+). Because each patient carries a sampling
    weight, a naive average would be biased toward over-sampled groups. Survey
    Means is the right tool because it incorporates the survey weights to
    produce design-correct estimates of the mean systolic blood pressure along
    with proper standard errors, and lets us compare those weighted means across
    the three age groups.
  >> VARIABLE SELECTION:
    - Analysis variable (mean): systolic_bp
    - Weight: weight
    - By / Domain: age_group (18-40, 41-65, 66+)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Systolic BP weighted mean = 129.09   SE = 0.79   95% CI (127.53 , 130.65)   CV 0.62%
    Stratified + weighted design (Taylor SE)

>> COMMENTARY (narration):
    National health surveys select patients not at random but with a stratified, weighted design. A simple mean would
    be wrong for such data; complex-sample analysis gives the correct mean and standard error (Taylor linearization)
    by accounting for the design weights and stratification. The systolic blood pressure weighted mean is 129.1, with
    a very narrow confidence interval (CV 0.62%). National health surveys (NHANES-like) are always designed this way;
    survey methods are essential for population-generalisable, unbiased estimation.

====================================================================================

#79  Survey Frequencies
    file: 79_survey_freq_smoking.xlsx
  >> SCENARIO (narration):
    Suppose we ran a population survey of 400 people and recorded their smoking
    status (standard, active low, active high, or none) along with their region
    (urban or rural), and each respondent has a sampling weight. We want
    weighted prevalence estimates of each smoking category rather than raw
    counts, and we'd like to see how the distribution differs between urban and
    rural areas. Survey Frequencies is appropriate because it applies the survey
    weights to estimate the population proportion in each smoking category with
    correct standard errors, optionally cross-tabulated by region.
  >> VARIABLE SELECTION:
    - Categorical variable (frequencies): smoking (standard, active low, active high, none)
    - Weight: weight
    - By / Domain: region (urban, rural)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    "active high" proportion = 0.042 (SE 0.010)   |   "active low" proportion = 0.185 (SE 0.020)
    Smoking status -- stratified + weighted design

>> COMMENTARY (narration):
    We estimated the proportion of smoking-status categories in the population, accounting for the complex sample
    design (with design-based standard errors). A simple percentage would give the wrong standard error by ignoring
    stratification and weights; survey frequency analysis produces a confidence interval reflecting the true
    uncertainty. It is the correct way to estimate risk-factor/behavior proportions in a population (smoking
    prevalence, vaccination rate, disease prevalence) in a design-consistent manner -- the basis of public-health
    monitoring.

====================================================================================

#80  Survey Totals
    file: 80_survey_total_diabetes_prediction.xlsx
  >> SCENARIO (narration):
    Suppose we surveyed 280 health institutions, sampled within three size
    strata (S-small, M-medium, L-large), and recorded the number of diabetes
    patients at each institution along with a sampling weight. We don't want an
    average here; we want to estimate the total number of diabetes patients
    across the whole population of institutions. Survey Totals is the correct
    procedure because it uses the survey weights to scale the observed counts up
    to a population total estimate with valid standard errors, and the strata
    let us see how that total breaks down by institution size.
  >> VARIABLE SELECTION:
    - Analysis variable (total): diabetes_patient_count
    - Weight: weight
    - Stratum / By: size_stratum (S-small, M-medium, L-large)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total diabetes patient count = 727,848   SE = 7,969   95% CI (712,160 , 743,536)
    Size-stratified design (Taylor SE)

>> COMMENTARY (narration):
    We estimated the total number of diabetes patients in a region/country using the sampling design and weights:
    about 728,000 (95% CI 712,000 - 744,000). Scaling from the sampled health institutions up to the whole population
    requires correct weighting. Survey total analysis does exactly this and expresses uncertainty with the standard
    error. It is the official public-health statistics method used to estimate population totals (total patients,
    total cases, total disease burden) from a sample -- critical for resource planning.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg_bp.xlsx
  >> SCENARIO (narration):
    Imagine we ran a multi-region cohort to understand what drives systolic
    blood pressure in hospitalized infection patients, but our sample wasn't a
    simple random draw. Patients were sampled with unequal probabilities across
    urban, rural and Marmara regions, so each record carries a survey weight. We
    want to model systolic_bp as a function of age and BMI while properly
    accounting for that complex sampling design. We use survey regression
    because an ordinary regression would underestimate the standard errors and
    give biased coefficients when the data come from a weighted, regionally
    stratified survey.
  >> VARIABLE SELECTION:
    - Dependent variable: systolic_bp
    - Predictors: age, BMI
    - Predictor (categorical, 3 levels): region (urban, rural, Marmara)
    - Survey weight: weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.547   age coefficient = 0.473   p < .001 ***
    Systolic BP ~ age + BMI   (stratified + weighted)

>> COMMENTARY (narration):
    We modelled systolic blood pressure from age and BMI, accounting for the complex sample design (strata, weights).
    The age effect is strong and significant (coefficient 0.47, p < .001): blood pressure rises markedly with age --
    a known physiological relationship. The model explains 55% of the variance. Design-based standard errors are more
    accurate than the (often too optimistic) errors of simple OLS. When estimating relationships from national health
    survey data (where the sample is not random), survey regression is necessary for valid inference.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic_diabetes.xlsx
  >> SCENARIO (narration):
    Suppose we want to estimate how age, BMI and region relate to the odds of a
    diabetes comorbidity among infection inpatients, using data from a weighted
    national surveillance survey rather than a simple random sample. The outcome
    diabetes is binary (present versus absent), and each patient has a sampling
    weight reflecting their selection probability. We fit a survey logistic
    regression so the odds ratios for age and BMI are estimated correctly and
    the confidence intervals respect the urban versus rural design weighting.
  >> VARIABLE SELECTION:
    - Dependent variable (binary): diabetes
    - Predictors: age, BMI
    - Predictor (categorical, 2 levels): region (urban, rural)
    - Survey weight: weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    age: OR = 1.051   p < .001 ***   95% CI (1.030 , 1.072)
    Diabetes ~ age + BMI   (stratified + weighted)

>> COMMENTARY (narration):
    Predicting a person's probability of having diabetes from age and BMI, we handled both the binary outcome
    (logistic) and the complex sample design together. The odds ratio for age is 1.051 (p < .001): for each one-year
    increase in age the odds of diabetes rise about 5% -- cumulatively a large difference (10 years = ~63% rise).
    Because the design weights and stratification are accounted for, the OR and confidence interval are
    design-consistent and valid. Survey logistic regression is the correct way to model binary outcomes (disease
    present/absent) from complex health surveys -- the standard of epidemiological risk estimation.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_mortality_age.xlsx
  >> SCENARIO (narration):
    Here we suspect that the effect of patient age on one-year mortality risk
    after a severe infection is not a straight line. Mortality probability may
    stay flat through middle age and then rise sharply in the elderly. We have
    age and the recorded 1-year mortality probability for each patient. We use a
    Generalized Additive Model because it lets the relationship between age and
    mortality bend through a smooth spline instead of forcing a linear slope,
    revealing the true clinical shape of the age-mortality curve.
  >> VARIABLE SELECTION:
    - Dependent variable: 1yr_mortality_probability
    - Smooth predictor: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (explained) = 0.983
    1-year mortality probability ~ s(age)  (smooth function)

>> COMMENTARY (narration):
    GAM is a flexible generalisation of linear regression: it models a variable's effect with "smooth curves"
    (splines) rather than a straight line. Age's effect on mortality probability is often nonlinear -- it rises fast
    beyond a certain age (super-exponential). GAM learns this curved pattern from the data without assuming its shape
    in advance; the model explains 98% of the variance (very high). It preserves interpretability while adding
    flexibility: we can read from the curve at which age mortality risk accelerates. It is the ideal method for
    modelling nonlinear age-/dose-related risk relationships in the clinic.

====================================================================================

#84  Discriminant Analysis (LDA/QDA)
    file: 84_diskriminant_3sinif_5lab.xlsx
  >> SCENARIO (narration):
    Imagine we want to classify infection patients into three severity classes
    -- culture_negative, mild and severe -- using a panel of routine labs
    measured on admission: white cell count, CRP, creatinine, AST and
    hemoglobin. The goal is to find the linear combinations of these labs that
    best separate the three groups and to build a rule that assigns new patients
    to a class. We use discriminant analysis because the predictors are
    continuous, the outcome is a small set of clinical classes, and it gives
    interpretable discriminant functions for severity triage.
  >> VARIABLE SELECTION:
    - Grouping variable (3 classes): class (culture_negative, mild, severe)
    - Predictors: WBC_K_uL, CRP_mg_L, creatinine, AST_U_L, Hb_g_dL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (3 disease classes)
    Separation from 5 lab variables (WBC, CRP, creatinine, AST, Hb)

>> COMMENTARY (narration):
    Discriminant analysis finds the linear combinations of variables that best separate the groups (three disease
    classes) -- a bit like a classification-focused PCA. Accuracy is 100%: five lab variables separate the three
    disease classes perfectly. Discriminant analysis both classifies and answers "which lab value matters most for
    the separation" (discriminant function loadings). It resembles logistic regression but is the classic choice for
    multiple groups when variables are normally distributed. In the clinic it is used in diagnostic classification
    and for assigning new patients to existing diagnostic groups.

====================================================================================

#85  Conditional Logit (Choice Model)
    file: 85_conditional_logit_treatment.xlsx
  >> SCENARIO (narration):
    Suppose we ran a discrete choice experiment to learn how infection patients
    trade off treatment duration against cost when picking a therapy. Each
    patient faced a choice set of four alternatives -- surgical, chemotherapy,
    radiotherapy and a targeted-oriented option -- each described by its
    duration in months and its cost in lira, and we recorded which one they
    chose. We use a conditional logit model because the choice depends on the
    attributes of the alternatives within each patient's choice set, letting us
    estimate how sensitive treatment choice is to duration and cost.
  >> VARIABLE SELECTION:
    - Choice set / case identifier: choice_id
    - Chosen indicator: chosen
    - Alternative attributes: duration_month, cost_TL
    - Alternative (categorical): treatment (T1-surgical, T2-chemotherapy, T3-radiotherapy, T4-targeted-oriented)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (McFadden) = 0.210
    Treatment choice ~ treatment duration (months) + cost (TL)

>> COMMENTARY (narration):
    We examined which treatment patients choose with a conditional logit model -- each patient selects one from a
    choice set. McFadden pseudo R2 = 0.21 is good for choice models (0.2-0.4 is strong). The model shows patients
    prefer shorter, lower-cost treatments. This method underlies "discrete choice" analysis in health economics and
    patient-preference research; it is powerful for modelling treatment preferences, patient-centred outcomes and
    cost/access effects in health policy.

====================================================================================

#86  Kaplan-Meier Survival Analysis
    file: 86_km_infection_survival.xlsx
  >> SCENARIO (narration):
    Imagine we followed 200 patients after diagnosis of a serious infection and
    recorded how many months each survived and whether they died or were still
    alive at last contact. We also know each patient's disease stage -- early,
    medium or advanced. We want to estimate and visualize the survival curves
    and compare them across stages. We use Kaplan-Meier analysis because it
    handles censored follow-up correctly and produces stage-specific survival
    curves we can compare with a log-rank test.
  >> VARIABLE SELECTION:
    - Time variable: duration_month
    - Event indicator: death
    - Stratification / group: stage (early, medium, advanced)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 41.1 months   (grouped by stage)
    Log-rank test compares groups

>> COMMENTARY (narration):
    We analysed patients' survival time -- by disease stage -- with Kaplan-Meier. Median survival is 41.1 months.
    Kaplan-Meier analyses "time-to-event" data -- correctly using censored patients who have not yet had the event
    (death) too -- and tests stage differences in survival with the log-rank test. In oncology and infectious
    diseases it is the basic visual and test for "time-event" data such as survival, recurrence, recovery; it is at
    the centre of prognosis and treatment-efficacy evaluation.

====================================================================================

#87  Cox Proportional Hazards Regression
    file: 87_cox_death_hazard.xlsx
  >> SCENARIO (narration):
    Here we want to know which factors raise the risk of death over follow-up in
    our infection cohort, adjusting for several variables at once. For each
    patient we have age, disease stage, whether they received the new treatment,
    a biomarker level, the follow-up time in months and whether death occurred.
    We use Cox proportional hazards regression because it models the hazard of
    death as a function of these covariates without assuming a particular
    baseline survival shape, giving adjusted hazard ratios for each predictor.
  >> VARIABLE SELECTION:
    - Time variable: duration_month
    - Event indicator: death
    - Covariates: age, stage, new_treatment, biomarker

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 134 (53.6%)   Concordance = 0.595
    age: HR = 1.013 p = 0.063 (borderline)   |   stage, new_treatment, biomarker: non-significant

>> COMMENTARY (narration):
    We modelled patients' death risk from clinical variables (age, stage, new treatment, biomarker) with Cox
    regression. In this model no predictor was strongly significant -- age is borderline (HR = 1.013, p = 0.063), the
    others non-significant; concordance 0.59 indicates modest discrimination. This is informative too: in this data
    death risk is not well explained by these chosen variables -- other factors (comorbidity, treatment response) are
    probably determining. Cox regression is the gold-standard survival model relating "time-to-event" to predictors;
    a non-significant result also shows which factors do not determine prognosis.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT) Model
    file: 88_aft_weibull_diyaliz.xlsx
  >> SCENARIO (narration):
    Suppose we are studying survival time in dialysis patients with infection
    and we care directly about how long they survive rather than just the
    hazard. We recorded follow-up time in months, whether death occurred, and
    whether the patient had diabetes. We fit a parametric accelerated failure
    time model with a Weibull distribution because it models survival time on a
    log scale and tells us, as a time ratio, how much diabetes accelerates or
    decelerates time to death.
  >> VARIABLE SELECTION:
    - Time variable: duration_month
    - Event indicator: death
    - Covariate: diabetes

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution = Weibull   Median survival = 40.30 months
    Survival ~ diabetes

>> COMMENTARY (narration):
    We analysed patients' survival time with a parametric Weibull model. While Cox is semi-parametric, Weibull models
    the whole survival curve with a specific mathematical distribution -- if the data fit it, it gives more efficient
    estimates and allows extrapolation. Median survival is 40.3 months. The AFT (accelerated failure time)
    interpretation is intuitive: how many times a factor like diabetes lengthens/shortens the survival time. When the
    time distribution is known, parametric models give more powerful and predictable results than Cox.

====================================================================================

#89  Competing Risks
    file: 89_competing_risks_3neden.xlsx
  >> SCENARIO (narration):
    Imagine our infection patients can die from more than one cause, and we want
    the cumulative incidence of each cause rather than treating all deaths the
    same. We recorded follow-up time in months, an event indicator, the specific
    cause of death code, plus age and disease stage. We use a competing risks
    analysis because a standard Kaplan-Meier would overestimate the probability
    of a given cause when other causes of death compete; this approach gives
    proper cause-specific cumulative incidence functions adjusted for age and
    stage.
  >> VARIABLE SELECTION:
    - Time variable: duration_month
    - Event status indicator: event
    - Competing event / cause code: death_cause
    - Covariates: age, stage

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 63   |   3 different causes of death (death_cause types)
    Separate cumulative incidence (CIF) and risk for each event type

>> COMMENTARY (narration):
    A patient can die from more than one cause: disease-related, treatment complication, or unrelated causes -- and
    these are mutually exclusive (once one happens the others cannot). Classic survival handles this incorrectly;
    competing-risks analysis produces a separate cumulative incidence for each cause of death. In the data 63 patients
    are still alive (censored), the rest died of different causes. This framework correctly answers questions like
    "which factor increases the risk of which cause of death". In medical survival research, when there are different
    causes of death, competing-risks analysis is essential -- otherwise one cause's risk is overstated.

====================================================================================

#90  Time-Dependent Cox Model
    file: 90_tvcox_procalcitonin_death.xlsx
  >> SCENARIO (narration):
    Here we have a biomarker, procalcitonin, that is measured repeatedly over
    time, so a patient's risk of death changes as their procalcitonin rises or
    falls. The data are in counting-process form with a start and end time for
    each interval, the procalcitonin value during that interval, and an event
    indicator for whether death occurred. We use a time-dependent Cox model
    because procalcitonin is a time-varying covariate, and this model lets the
    hazard of death respond to the patient's current procalcitonin level rather
    than a single baseline measurement.
  >> VARIABLE SELECTION:
    - Interval start time: start
    - Interval stop time: end
    - Event indicator: event
    - Time-dependent covariate: procalcitonin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    procalcitonin: HR = 1.018   95% CI (0.956 , 1.083)   p = 0.576 (non-significant)
    Time-varying procalcitonin covariate

>> COMMENTARY (narration):
    A patient's procalcitonin changes over time; standard Cox assumes it constant, time-dependent Cox uses the
    current value in each time interval (start-stop format). In this model the effect of current procalcitonin on
    death risk is non-significant (HR = 1.018, p = 0.576) -- in this data the momentary procalcitonin change does not
    predict death timing. A non-significant result is information too. Time-dependent Cox is the correct method when
    "the current, continuously monitored condition matters, not just the baseline"; it is the standard for modelling
    time-varying biomarkers in longitudinal clinical monitoring.

====================================================================================

#91  Survey-Weighted Cox Regression
    file: 91_survey_phreg_death.xlsx
  >> SCENARIO (narration):
    Imagine we want to know which factors drive mortality among hospitalized
    infection patients, but our patients were not sampled uniformly. They were
    drawn from clinics across three regions (A, B, and C) using a complex survey
    design, so each patient carries a sampling weight. We followed each patient
    over time, recording duration_month until death or censoring, and we want to
    estimate how age and region affect the hazard of death while correctly
    accounting for the survey weights. A survey-weighted Cox regression is the
    right tool here because it models time-to-event data and produces
    population-level hazard ratios that respect the unequal selection
    probabilities and clinical clustering.
  >> VARIABLE SELECTION:
    - Time / Duration: duration_month
    - Event (1=death, 0=censored): death
    - Covariate (continuous): age
    - Covariate (categorical, 3 levels): region (A, B, C)
    - Weight: weight
    - Cluster / Strata: clinical_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.523   age: HR = 1.013, p = 0.006 **
    Death ~ age   (weighted/clustered, centre cluster)

>> COMMENTARY (narration):
    We combined survival analysis with a complex sample design: patients come from a population sampled with weights
    and clustered by centres. Survey-PHREG runs the Cox model with design weights and cluster-robust standard errors,
    so the hazard ratios and confidence intervals generalise to the population. Age is significant (HR = 1.013, p =
    0.006) -- death risk rises with age. Concordance 0.52 means age alone is a weak discriminator. When national
    health records have "time-to-event" data, ignoring the design biases the estimates; survey-PHREG provides valid
    inference.

====================================================================================

#92  Interval-Censored Survival Analysis
    file: 92_interval_censored_nuks.xlsx
  >> SCENARIO (narration):
    Suppose we are studying recurrence of infection after treatment, but we only
    see patients at scheduled follow-up visits, so we never observe the exact
    recurrence day. For each of these 60 patients we know only an interval: the
    last visit when they were still recurrence-free, left_censor, and the visit
    by which recurrence had occurred, survival_censor. We also recorded disease
    stage as early, medium, or advanced. Because the event time is bracketed
    between two observation points rather than known exactly, an interval-
    censored survival analysis is appropriate, letting us estimate the survival
    curve and compare recurrence timing across stages without biasing the
    estimates.
  >> VARIABLE SELECTION:
    - Left interval bound (last recurrence-free time): left_censor
    - Right interval bound (recurrence observed by): survival_censor
    - Group / Stratum (3 levels): stage (early, medium, advanced)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 30   Median survival = 10.0 months   (Turnbull algorithm)
    Recurrence time known within [lower, upper] interval

>> COMMENTARY (narration):
    In clinical monitoring we rarely know the exact time of an event (recurrence): we call the patient for periodic
    checks, find them clear at one and with recurrence at the next -- the recurrence happened somewhere between.
    Interval-censored survival (Turnbull algorithm) handles exactly this uncertainty, avoiding fixing the event to an
    arbitrary date. Median time to recurrence is 10 months. By the nature of periodic checks (every 3 months, annual
    screening), event times are always within an interval; forcing this data to a single point biases the result.
    The interval-censored method matches the reality of periodic monitoring data such as cancer recurrence follow-up.

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty_cox_clinical.xlsx
  >> SCENARIO (narration):
    Imagine we are following 120 infection patients treated across several
    clinics, and we suspect that patients within the same clinic share
    unmeasured risk because of local protocols, staffing, or case mix. We
    recorded each patient's age, their follow-up duration_month, and whether
    they died. We want to estimate how age affects the hazard of death while
    allowing each clinic to have its own baseline risk level. A frailty Cox
    model fits this perfectly: it adds a random clinic-level frailty term to the
    proportional hazards model, capturing the within-clinic correlation that an
    ordinary Cox model would ignore.
  >> VARIABLE SELECTION:
    - Time / Duration: duration_month
    - Event (1=death, 0=censored): death
    - Covariate (continuous): age
    - Frailty / Cluster (random group): clinical_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.513   Centre-clustered shared frailty (random effect)
    Death ~ age  +  (centre frailty)

>> COMMENTARY (narration):
    Patients are clustered within centres/clinics; patients in the same centre share unmeasured common factors
    (treatment protocol, experience, resources) and carry similar risk. Frailty Cox adds a shared "frailty" (random
    effect) for each centre to model this clustering -- the mixed-model version of survival analysis. Concordance
    0.51. Accounting for centre-level hidden differences gives both correct standard errors and information on "how
    much heterogeneity exists between centres". It is the correct method for multi-centre clinical survival data
    (hospitals, clinics).

====================================================================================

#94  Time Series Description
    file: 94_ts_monthly_case.xlsx
  >> SCENARIO (narration):
    Suppose the infection control team wants a first look at how the monthly
    number of new cases has behaved over the past five years. We have 60
    consecutive monthly observations: a date for each month and the monthly_case
    count. Before fitting any forecasting model, we simply want to describe the
    series, its overall level, the spread, whether the counts are trending up or
    down, and whether there is an obvious seasonal rhythm. A time series
    description is the natural starting point because it summarizes and
    visualizes the sequence over time, giving us the context we need before any
    modeling.
  >> VARIABLE SELECTION:
    - Time / Date index: date
    - Series / Value: monthly_case

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.998 -> NOT stationary   Trend = increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly case-count series. The ADF test does not find it stationary (p = 0.998) -- the mean changes
    over time, there is a trend (increasing) and seasonality is present (seasonal fluctuation of infections). This
    pre-diagnosis is critical: most time-series models (like ARIMA) require stationarity, so differencing/
    transformation may be needed first. Time-series analysis reveals the structure of long surveillance data --
    trend, seasonality, stationarity -- and tells which modelling steps are required. It is the basic analysis of
    epidemiological surveillance (case counts, incidence).

====================================================================================

#95  STL Decomposition
    file: 95_stl_YBU_daily.xlsx
  >> SCENARIO (narration):
    Imagine we are monitoring intensive care unit pressure during an outbreak
    and we have three full years of daily ICU bed occupancy, 1095 days in total,
    each with a date and a YBU_occupancy value. We can see the numbers wobble
    day to day, but we suspect there is both a slow underlying trend and a
    repeating weekly or seasonal pattern. To untangle these, we apply STL
    decomposition, which splits the daily occupancy series into trend, seasonal,
    and remainder components. This lets us judge whether occupancy is genuinely
    rising over time or just reflecting predictable seasonal cycles.
  >> VARIABLE SELECTION:
    - Time / Date index: date
    - Series / Value: YBU_occupancy

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Period (seasonal length) = 7 (daily -> weekly cycle)
    Series decomposed into trend + seasonal + residual components

>> COMMENTARY (narration):
    STL (Seasonal-Trend decomposition using Loess) splits a daily ICU-occupancy series into three components:
    long-term trend, a recurring cycle (7-day = weekly) and the remaining residual. The weekly period is meaningful
    -- ICU occupancy fluctuates regularly across the days of the week (e.g. weekends differ). This decomposition
    clarifies "is occupancy generally rising, or just fluctuating weekly". STL is the most intuitive way to interpret
    seasonal/cyclical health series, separating trend from cycle so each can be assessed separately.

====================================================================================

#96  ARIMA Forecasting
    file: 96_arima_monthly_ddd.xlsx
  >> SCENARIO (narration):
    Suppose the antimicrobial stewardship program wants to forecast antibiotic
    consumption so the pharmacy can plan stock. We have 120 months of defined
    daily doses, a date for each month and the monthly_ddd value. The series
    shows persistence from month to month and some seasonal movement, and we
    want a principled short-term forecast with confidence intervals. ARIMA
    forecasting is well suited here because it models the autocorrelation and
    any differencing needed for stationarity, producing forward projections of
    antibiotic use along with uncertainty bands.
  >> VARIABLE SELECTION:
    - Time / Date index: date
    - Series / Value to forecast: monthly_ddd

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model fitted   AIC = 2471.69
    Monthly antibiotic consumption (DDD) -- future forecast

>> COMMENTARY (narration):
    We fitted an ARIMA model to a monthly antibiotic-consumption (DDD -- defined daily dose) series and forecast the
    future. ARIMA combines the series' own past values (AR), past errors (MA) and differencing (I, to make it
    stationary). With AIC = 2472 the best model was selected. Forecasts rest on the observed trend and
    autocorrelation structure and come with an uncertainty band. Antibiotic-consumption forecasting matters for
    antibiotic stewardship and resistance control: anticipating future consumption enables intervention and stock
    planning. ARIMA is the classic forecasting method for health series with seasonality and trend.

====================================================================================

#97  Exponential Smoothing (Holt-Winters)
    file: 97_ets_monthly_emergency.xlsx
  >> SCENARIO (narration):
    Imagine the emergency department wants to anticipate patient volume so it
    can staff appropriately. We have 84 months of emergency_application counts,
    each tied to a date, and the series clearly carries both a gradual trend and
    a yearly seasonal pattern of infection-driven visits. Exponential smoothing
    with the Holt-Winters method is ideal here because it explicitly models
    level, trend, and seasonality with smoothing parameters, giving a smooth,
    seasonally aware forecast of emergency applications for the coming months.
  >> VARIABLE SELECTION:
    - Time / Date index: date
    - Series / Value to forecast: emergency_application

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method = Holt-Winters (additive seasonal, period 12)   AIC = 1062.52
    Monthly emergency application count forecast

>> COMMENTARY (narration):
    Holt-Winters exponential smoothing estimates the level + trend + seasonality components in a weighted way (giving
    more weight to the recent past). It captured the seasonal pattern of monthly emergency-department applications
    (AIC = 1063). Emergency applications typically show strong seasonality (flu season, temperature); Holt-Winters is
    often very successful for such series and is more intuitive than ARIMA. In health-demand forecasting (emergency
    applications, bed needs) it is an easy-to-build, interpretable alternative; it is valuable for capacity and
    staffing planning.

====================================================================================

#98  Mann-Kendall Trend and Sen's Slope
    file: 98_mann_kendall_survival_expectation.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether survival prospects for a serious infection
    have genuinely improved over the decades rather than just fluctuating year
    to year. We have 40 yearly observations: the year and the
    survival_expectation for patients diagnosed that year. We are not assuming a
    straight line, we just want to detect a monotonic trend and quantify its
    magnitude. The Mann-Kendall trend test plus Sen's slope is the right choice:
    it is a nonparametric way to confirm whether survival expectation is
    trending upward over time and to estimate the per-year rate of change
    robustly.
  >> VARIABLE SELECTION:
    - Time index: year
    - Series / Value: survival_expectation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   Trend = INCREASING
    Life expectancy long-term trend

>> COMMENTARY (narration):
    We tested whether life expectancy shows a significant long-term trend with the Mann-Kendall test -- a
    nonparametric, outlier- and distribution-robust trend test. The result is highly significant (p < .001) and the
    direction is increasing: life expectancy is rising consistently over the years. This may reflect the positive
    cumulative effect of health services/treatments. Mann-Kendall + Sen's slope is the gold standard for detecting
    long-term health trends -- from demography to epidemiology; Sen's slope gives the trend magnitude without being
    affected by outliers.

====================================================================================

#99  Anomaly Detection
    file: 99_anomali_lab_results.xlsx
  >> SCENARIO (narration):
    Imagine the lab quality team wants to automatically flag unusual results
    that may signal data-entry errors or genuinely critical infection cases. For
    200 patients we have a panel of routine lab values: WBC_K_uL, Hb_g_dL,
    creatinine, AST_U_L, and CRP_mg_L. We want to identify patients whose
    combined profile is far from the typical pattern. Anomaly detection is
    appropriate because it scores each patient on how unusual their multivariate
    lab profile is, surfacing outliers that deserve a second look without us
    having to set arbitrary single-test cutoffs.
  >> VARIABLE SELECTION:
    - Feature: WBC_K_uL
    - Feature: Hb_g_dL
    - Feature: creatinine
    - Feature: AST_U_L
    - Feature: CRP_mg_L

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 17   (IQR-based multivariate outlier detection)
    Extraordinary profiles in lab results

>> COMMENTARY (narration):
    We found the extraordinary profiles in patients' lab results (WBC, Hb, creatinine, AST, CRP) with anomaly
    detection -- 17 patients were flagged as deviating from the normal pattern. These anomalies can point to real
    clinical situations: a rare disease, a severe complication, or a laboratory/preanalytical error. Anomaly
    detection automatically picks out the few critical cases that need attention from large laboratory databases. In
    the clinic it is extremely practical for early warning, critical-value notification and laboratory quality
    control.

====================================================================================

#100  Variance Components Analysis
    file: 100_varcomp_clinical_doctor_patient_h2.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand where the variability in treatment response
    really comes from in our infection patients. The same outcome,
    treatment_response, was measured for 120 patients nested within doctors
    (Kli_Dr01 through Kli_Dr05), who are themselves nested within four clinics
    (clinical_A to clinical_D). We suspect some clinics and some doctors are
    systematically more effective than others. A variance components analysis is
    the right approach because it partitions the total variance in treatment
    response into clinic-level, doctor-level, and patient-level pieces, telling
    us how much of the differences in outcome is attributable to each layer of
    the hierarchy.
  >> VARIABLE SELECTION:
    - Response / Dependent: treatment_response
    - Random factor (top level): clinical (clinical_A, clinical_B, clinical_C, clinical_D)
    - Random factor (nested): doctor_id (Kli_Dr01...Kli_Dr05)
    - Random factor (unit level): patient_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    clinical = 40.1%   doctor (nested in clinical) = 13.4%   residual = 46.5%
    Nested: clinical > doctor > patient (REML)

>> COMMENTARY (narration):
    We partitioned the variability of treatment response by which level -- clinical, doctor or patient -- it comes
    from with variance-components analysis: clinical and doctor are nested. The result: 40.1% of the variance comes
    from between-clinic differences, 13.4% from between-doctor (within-clinic) differences, and 46.5% individual/
    residual variation. So the largest structural source is the clinic level -- clinics differ markedly in treatment
    response (protocol, resource differences). This is critical for quality improvement: it shows where the
    variability concentrates and at which level (clinic or doctor) an intervention should be targeted.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old_treatment.xlsx
  >> SCENARIO (narration):
    Suppose we are testing whether a newly developed antibiotic regimen actually
    improves clinical response compared to the standard-of-care regimen in
    patients with bacterial infection. Seventy patients were assigned to either
    the standard or the new treatment, and after the course we recorded a
    composite response_score for each one. Instead of just asking whether the
    difference is significant, we want to quantify the evidence directly. We use
    a Bayesian t-test because it gives us a Bayes factor telling us how much the
    data favor the new treatment over no difference, and a posterior on the mean
    difference rather than a single p-value.
  >> VARIABLE SELECTION:
    - Group (2 levels): treatment (standard, new)
    - Dependent variable: response_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 2.81e+40 (decisive evidence)   Cohen's d = 3.94
    New treatment vs old treatment -- response score

>> COMMENTARY (narration):
    We tested whether a new treatment differs from the old with a Bayesian t-test. The Bayes Factor BF10 = 2.81e+40 --
    astronomically large, "decisive evidence": the data support the hypothesis of a difference quadrillions of times
    more than no difference. Cohen's d = 3.94 makes the effect enormous. Unlike a classic p-value, the Bayes Factor
    directly measures the strength of evidence for both H1 and H0 and distinguishes "absence of evidence" from
    "evidence of absence". The new treatment's effect is not just statistical but clinically overwhelming -- the
    Bayesian framework shows this convincingly.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BMI_procalcitonin_BF10.xlsx
  >> SCENARIO (narration):
    Here we want to know whether a patient's body mass index is associated with
    their procalcitonin level, since obesity is often linked to a heightened
    inflammatory state in infection. We have 80 patients with both BMI and
    procalcitonin measured. Rather than report only a correlation coefficient
    and a p-value, we want a Bayes factor that tells us how strongly the data
    support a real correlation versus no association. We use Bayesian
    correlation because it gives us the BF10 and a posterior credible interval
    for the correlation, which is far more informative for a small clinical
    sample.
  >> VARIABLE SELECTION:
    - Variable 1: BMI
    - Variable 2: procalcitonin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.662 (strong)   BF10 = 4.77e+08 (decisive evidence)
    BMI - procalcitonin relationship

>> COMMENTARY (narration):
    We assessed the relationship between BMI and procalcitonin with Bayesian correlation: r = 0.66 is strong and the
    Bayes Factor BF10 = 4.77e+08 makes the evidence for the relationship's existence decisive. As BMI rises,
    procalcitonin (an inflammation marker) increases -- a sign of the known link between obesity and low-grade chronic
    inflammation. Unlike classic correlation, Bayesian correlation expresses the strength of the relationship with a
    probability distribution and an evidence factor -- also answering "how sure are we". It is a robust framework for
    quantifying strong links between clinical variables.

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_antibiotic_sex.xlsx
  >> SCENARIO (narration):
    In this study we are comparing three antibiotics, labeled A, B and C, to see
    whether they differ in how much they reduce procalcitonin, and whether the
    effect also depends on patient sex. We measured procalcitonin in 120
    patients across the three drug groups and both sexes. We want to weigh the
    evidence for main effects and the interaction directly rather than running a
    stack of pairwise tests. We use a Bayesian ANOVA because it compares
    competing models with Bayes factors, telling us which factors the data
    actually support as drivers of procalcitonin.
  >> VARIABLE SELECTION:
    - Factor 1 (3 levels): antibiotic (A, B, C)
    - Factor 2 (2 levels): sex (female, male)
    - Dependent variable: procalcitonin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    antibiotic: BF10 = 2.59e+08 (decisive evidence)
    Antibiotic x sex -> procalcitonin

>> COMMENTARY (narration):
    We examined the effect of antibiotic and sex factors on procalcitonin with Bayesian ANOVA. The antibiotic effect
    is overwhelming: Bayes Factor 2.59e+08 -- "decisive evidence". So different antibiotics produce markedly different
    effects on procalcitonin (the infection response). Bayesian ANOVA gives a separate Bayes Factor for each effect
    and interaction, answering "which factor really matters" more informatively than classic ANOVA -- and can even
    provide "evidence of absence" for non-significant effects. It is a powerful way to evaluate effect strength with
    evidence in clinical experiments (drug x demographic factor).

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_clinical_random.xlsx
  >> SCENARIO (narration):
    We collected procalcitonin levels from 200 infection patients treated across
    five different clinics in Ankara, Istanbul, Izmir, Bursa and Antalya, and we
    also recorded each patient's age and BMI. Patients within the same clinic
    are not independent, since each center has its own case mix and treatment
    habits. We want to estimate how age and BMI relate to procalcitonin while
    letting each clinic have its own baseline level. We use a Bayesian
    hierarchical model because it partially pools information across clinics,
    giving stable clinic-level estimates and full posterior uncertainty for the
    fixed effects.
  >> VARIABLE SELECTION:
    - Dependent variable: procalcitonin
    - Grouping / random effect: clinical (K_Ankara, K_Istanbul, K_Izmir, K_Bursa, K_Antalya)
    - Predictors: age, BMI

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group (clinic) random variance = 14.42 (61.7%)   residual = 8.97 (38.3%)   ICC = 0.617
    Procalcitonin ~ age + BMI  +  (1 | clinic)   (Bayesian)

>> COMMENTARY (narration):
    We analysed multi-centre data (patients within clinics) with a Bayesian hierarchical model -- the Bayesian
    counterpart of LMM. The clinic-level variance is 61.7% of the total (ICC = 0.617); most of the variability comes
    from between-clinic differences. So procalcitonin level depends strongly on which clinic one is in (protocol/
    population difference). The Bayesian hierarchical model's strength: it expresses both within- and between-group
    uncertainty with full probability distributions and balances clinics with few observations via "partial pooling".
    It is a modern method for multi-centre clinical studies.

====================================================================================

#105  Spatial Lag (SAR) Model
    file: 105_spatial_sar_COVID_case_ratio.xlsx
  >> SCENARIO (narration):
    We are mapping COVID-19 burden across 100 provinces, where case_ratio is the
    per-capita case rate in each province. Epidemics spread between neighboring
    regions, so the outbreak in one province tends to raise the rate in adjacent
    ones. We want to know how socioeconomic level and mean population age relate
    to the case ratio while accounting for this spillover. We use a Spatial Lag
    (SAR) model because it explicitly includes a spatially lagged dependent
    variable, capturing the contagion-like diffusion of cases between
    neighboring provinces.
  >> VARIABLE SELECTION:
    - Dependent variable: case_ratio
    - Predictors: socioeconomic, age_mean
    - Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho (spatial lag) = 0.32   z = 3.95   p < .001   Pseudo R2 = 0.77
    Case ratio ~ socioeconomic + mean age  +  neighbour case ratio

>> COMMENTARY (narration):
    Modelling the COVID case ratio by province/region, we accounted for spatial dependence -- the tendency of nearby
    regions to be similar. The SAR (spatial autoregressive) model links a region's case ratio to its neighbours'
    ratio too (the rho parameter). Here rho = 0.32, significantly positive (z = 3.95, p < .001): the neighbour effect
    is real and moderate -- a region's case ratio is influenced by neighbouring regions (transmission spread). The
    model explains 77% of the variance. Using a spatial model is critical, because ordinary regression gives spurious
    significant results when spatial autocorrelation is present. SAR is the right method for modelling spatial
    transmission processes such as outbreak spread.

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_disease_residual_autocorr.xlsx
  >> SCENARIO (narration):
    Again working with COVID-19 case ratios across 100 provinces, we suspect
    that important drivers of the outbreak, such as mobility or unmeasured local
    policy, are themselves spatially clustered and end up in the model
    residuals. If we ignore that, our standard errors for socioeconomic level
    and mean age will be wrong. We want clean, unbiased estimates of these
    predictors. We use a Spatial Error model because it places the spatial
    dependence in the error term, correcting for autocorrelated residuals rather
    than modeling direct case-to-case spillover.
  >> VARIABLE SELECTION:
    - Dependent variable: case_ratio
    - Predictors: socioeconomic, age_mean
    - Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda (spatial error) = 0.63   z = 7.08   p < .001   Pseudo R2 = 0.72
    Case ratio ~ variables  +  spatial error structure

>> COMMENTARY (narration):
    The spatial error model addresses a different kind of spatial dependence than SAR: not a direct effect of
    neighbour values, but correlation induced through the residuals by unobserved (omitted) spatial factors. Lambda =
    0.63 (z = 7.08, p < .001) shows this spatial error structure is strong -- there are unmeasured common spatial
    factors (health infrastructure, mobility, policy). Which spatial model is appropriate (SAR or SEM) depends on
    whether the effect comes from neighbour values or from unmeasured factors. SEM cleans out these unmeasured
    spatial confounders to estimate the true effect of the variables more accurately.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local_risk_factor.xlsx
  >> SCENARIO (narration):
    We want to understand whether the drivers of COVID-19 case rates are the
    same everywhere or whether they shift from region to region across 110
    provinces. For example, socioeconomic level might matter strongly in some
    areas but barely at all in others. A single global regression would hide
    this. We use Geographically Weighted Regression because it fits a local
    model around each province, producing maps of how the effects of
    socioeconomic level and mean age on case_ratio vary across space.
  >> VARIABLE SELECTION:
    - Dependent variable: case_ratio
    - Predictors: socioeconomic, age_mean
    - Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.811   (local/location-specific regression)
    Spatial variation of the risk-factor-case relationship

>> COMMENTARY (narration):
    Ordinary regression assumes a single risk-factor-case relationship for the whole region; but this relationship
    can vary in space -- a risk factor's effect may be strong in one region and weak in another. GWR (Geographically
    Weighted Regression) captures this spatial heterogeneity by fitting a separate local regression at each location
    (R2 = 0.81). The output is a map of coefficients: a surface showing where a risk factor's effect is strong and
    where it is weak. This tests "is the relationship the same everywhere" and sets local health-policy priorities.
    In outbreak and disease risk mapping, GWR reveals what the global model hides.

====================================================================================

#108  Nested Mixed Model (Nested LMM)
    file: 108_nested_lmm_clinical_doctor_visit_procalcitonin.xlsx
  >> SCENARIO (narration):
    In this longitudinal infection study, 288 patients were followed over four
    visits, V1 through V4, while procalcitonin was tracked. Each patient is
    treated by one doctor, and each doctor works within one clinic in Ankara,
    Istanbul, Izmir or Bursa, so doctors are nested inside clinics and patients
    inside doctors. We want to model how procalcitonin changes over visits while
    respecting this strict hierarchy. We use a Nested Mixed Model because the
    random effects are properly nested, letting us separate clinic-level,
    doctor-level and patient-level variability in the response.
  >> VARIABLE SELECTION:
    - Dependent variable: procalcitonin
    - Fixed effect: visit (V1, V2, V3, V4)
    - Nested random effects: clinical (K_Ankara, K_Istanbul, K_Izmir, K_Bursa) / doctor_no (D01-D06)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    clinic (P): F = 10.37, p < .001 ***   |   visit (R): F = 14.20, p < .001 ***
    doctor [within clinic] (F): F = 7.56, p < .001 ***

>> COMMENTARY (narration):
    In this design the measurements are nested: clinic > doctor > visit. The nested mixed model tests each level's
    contribution to the variance separately. The results are significant at all levels: there are real differences
    between clinics (F = 10.37), between doctors within a clinic (F = 7.56), and between visits (F = 14.20). Mixing up
    the levels in nested data (e.g. ignoring the clinic) produces spurious significance or wrong standard errors.
    Nested LMM is the correct framework for hierarchical clinical designs (hospital/doctor/patient, clinic/doctor/
    visit), separating the genuine contribution of each scale.

====================================================================================

#109  Crossed Mixed Model (Crossed LMM)
    file: 109_crossed_lmm_antibiotic_clinical.xlsx
  >> SCENARIO (narration):
    Here we evaluate five antibiotics, antibiotic_1 through antibiotic_5, each
    tested across four clinics K_A to K_D, with procalcitonin as the outcome in
    160 observations. Crucially, every antibiotic is used in every clinic, so
    the two grouping factors cross rather than nest. We want the average effect
    of antibiotic on procalcitonin while accounting for clinic-to-clinic
    variability that applies to all drugs. We use a Crossed Mixed Model because
    it treats antibiotic and clinic as crossed random effects, correctly
    partitioning variance when the factors are fully crossed.
  >> VARIABLE SELECTION:
    - Dependent variable: procalcitonin
    - Crossed random effect 1: antibiotic (antibiotic_1, antibiotic_2, antibiotic_3, antibiotic_4, antibiotic_5)
    - Crossed random effect 2: clinical (K_A, K_B, K_C, K_D)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    antibiotic: F = 2.24, p = 0.125 (non-significant)   |   clinical: F = 4.31, p = 0.028 * (significant)
    Crossed design -> procalcitonin

>> COMMENTARY (narration):
    Unlike nesting, here two factors are crossed: each antibiotic was administered in each clinic (fully factorial).
    The crossed mixed model handles this. The antibiotic main effect is non-significant (F = 2.24, p = 0.125), the
    clinic effect significant (F = 4.31, p = 0.028): the main determinant of procalcitonin level is not which
    antibiotic was given but which clinic one is in -- between-clinic protocol/population differences dominate. This
    answers "which factor really matters when two are tested jointly". The crossed mixed model is the right analysis
    for clinical designs where factors are evaluated simultaneously (drug x centre).

====================================================================================

#110  KDE Density Map
    file: 110_KDE_hastane_density_haritasi.xlsx
  >> SCENARIO (narration):
    We have the locations of 95 reporting units, each with a recorded case_ratio
    reflecting local disease intensity. Rather than looking at isolated points,
    public health teams want a smooth picture of where cases concentrate so they
    can target resources. We use a Kernel Density Estimation map because it
    turns the scattered point locations into a continuous density surface,
    highlighting outbreak hotspots and quiet zones across the study area.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon
    - Weight / intensity: case_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    95 unit points   mean case ratio = 408.6
    Case/hospital-density map via kernel density estimation

>> COMMENTARY (narration):
    On the Map tab, we render the geographic density of cases/health units into a heat map with KDE (Kernel Density
    Estimation). KDE turns points into a smooth density surface: we can see where case clustering is dense and where
    there are gaps. The distribution of 95 units shows how the disease burden concentrates geographically. This is
    used both to detect high-burden regions (intervention priority) and surveillance gaps. It is the basic mapping
    tool for visualising spatial health data; in outbreak management it guides resource distribution.

====================================================================================

#111  Hexbin Map
    file: 111_Hexbin_case_distribution.xlsx
  >> SCENARIO (narration):
    During a regional outbreak we logged the geographic coordinates of every
    reported infection across the surveillance area, and we want to see where
    cases pile up rather than just plotting 150 overlapping dots on a map. For
    each reporting unit we have its latitude and longitude, plus a
    density_interval label that flags the location as low, medium, or high case
    density. A hexbin map is ideal here because it bins the point coordinates
    into equal-area hexagons and colors each hexagon by how many cases fall
    inside it, turning a cluttered point cloud into a clean heat surface that
    immediately reveals the high-burden zones where containment efforts should
    be concentrated.
  >> VARIABLE SELECTION:
    - Coordinates (Latitude): lat
    - Coordinates (Longitude): lon
    - Weight / Value (optional): density_interval (high, medium, low)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    150 case/unit points   aggregated into hexagonal cells
    Each cell shows the point density within it by colour

>> COMMENTARY (narration):
    The hexbin map aggregates many points into hexagon cells and shows each cell's density by colour -- a clear
    solution, alternative to KDE, when thousands of cases overlap into an unreadable mess. We rendered 150 case
    points into a hexagonal grid to visualise where case density is highest. Hexagonal cells carry less edge-bias
    than squares and represent neighbour relations more evenly. It is the practical way to turn dense spatial health
    data (case location, transmission point) into a readable density map.

====================================================================================

#112  Moran's I (Spatial Autocorrelation)
    file: 112_Morans_I_COVID_spatial_autocorrelation.xlsx
  >> SCENARIO (narration):
    We mapped the COVID-19 case ratio for 80 administrative units and we suspect
    the epidemic is not scattered randomly but spills over between neighboring
    areas. The question is whether units with high case ratios sit next to other
    high-ratio units, which would signal geographic clustering of transmission.
    Each unit carries its lat and lon centroid and a continuous case_ratio
    value. Moran's I is the right tool because it quantifies global spatial
    autocorrelation, testing whether the case_ratio is significantly clustered,
    dispersed, or random across space given the spatial weights built from the
    coordinates.
  >> VARIABLE SELECTION:
    - Coordinates (Latitude): lat
    - Coordinates (Longitude): lon
    - Analysis variable: case_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.684   z = 12.01   p < .001
    Spatial autocorrelation of COVID case ratio (very strong positive)

>> COMMENTARY (narration):
    Moran's I is the global spatial-autocorrelation statistic measuring "do nearby regions have similar case ratios".
    The result is very strongly positive (I = 0.68, z = 12.01, p < .001): the COVID case ratio is densely spatially
    clustered -- high-ratio regions lie near each other, and so do low ones. This is not chance; transmission, shared
    mobility and socio-economic environment make neighbouring regions similar. This finding is critical in outbreak
    epidemiology and methodologically important: if spatial autocorrelation exists, ordinary statistics are wrong
    (hence the spatial models in #105-107). Moran's I is the starting diagnostic of spatial analysis.

====================================================================================

#113  Getis-Ord Hotspot
    file: 113_Getis_Ord_disease_hotspot.xlsx
  >> SCENARIO (narration):
    Knowing that the COVID case ratios are spatially clustered is useful, but
    for resource allocation we need to know exactly which regions are
    statistically significant hot spots versus cold spots. We have 68 reporting
    units, each with its lat and lon location and an observed case_ratio. The
    Getis-Ord Gi* statistic is perfect for this: it scans each unit together
    with its neighbors and returns a local z-score telling us where high case
    ratios cluster intensely (hot spots) and where low ratios cluster (cold
    spots), so public health teams can target outbreak response to the genuinely
    significant hot zones.
  >> VARIABLE SELECTION:
    - Coordinates (Latitude): lat
    - Coordinates (Longitude): lon
    - Analysis variable: case_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    33 hot spots (high-case cluster)   30 cold spots (low-case cluster)   n = 68
    Getis-Ord Gi* local statistic (z > 1.96 significant)

>> COMMENTARY (narration):
    While Moran's I reports the overall clustering of the whole region, Getis-Ord Gi* shows WHERE the hot/cold spots
    are, location by location. 33 regions are statistically significant "hot spots" (high case ratio surrounded by
    high ratio -- outbreak clusters), and 30 regions are "cold spots" (low-case clusters). This gives direct guidance
    for outbreak management: it is most efficient to focus testing, vaccination and intervention resources on the
    hot-spot clusters. Gi* hot-spot analysis is the standard spatial method for finding "where the disease
    concentrates geographically" -- it is at the centre of outbreak mapping.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_case_kumeleri.xlsx
  >> SCENARIO (narration):
    We pulled the geocoded locations of 115 confirmed infection cases and we
    want to detect natural disease clusters without deciding in advance how many
    clusters exist or treating isolated cases as part of a group. Each case is
    recorded only by its lat and lon coordinates. DBSCAN is well suited here
    because it groups cases that are densely packed in space into clusters while
    labeling sparse, far-apart cases as noise, which lets us pinpoint genuine
    transmission clusters and distinguish them from scattered sporadic cases on
    the map.
  >> VARIABLE SELECTION:
    - Coordinates (Latitude): lat
    - Coordinates (Longitude): lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   noise (isolated) = 11   n = 115
    Case locations clustered by density

>> COMMENTARY (narration):
    On the map we clustered infection cases' geographic locations by density with DBSCAN: four dense case groups were
    found, and 11 cases were flagged as "isolated/noise" points belonging to no cluster. DBSCAN's strength is that it
    does not require the number of clusters in advance and can capture irregularly shaped clusters -- and
    automatically separates sparse/distant cases. The four clusters show cases gather at particular foci (a possible
    common source/transmission chain); the isolated cases may be sporadic. This spatial clustering is a practical tool
    for defining outbreak foci, guiding contact tracing and identifying intervention regions. With this we complete
    the full 114-analysis tour of the Infectious Diseases package.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_CRP_mg_L.xlsx
  >> SCENARIO (narration):
    We follow 40 patients measured at three day levels (day0, day3, day7); each belongs to one of two antibiotic groups (regimen_A / regimen_B). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on CRP.
  >> VARIABLE SELECTION:
    - Dependent variable: CRP_mg_L
    - Subject ID: patient_id
    - Between-subjects factor: antibiotic
    - Within-subjects factor: day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (antibiotic): F(1,38) = 27.45  p < .001  np2 = 0.419
    Within (day):  F(2,76) = 341.87  p < .001  np2 = 0.900
    Interaction:         F(2,76) = 7.42  p = 0.001  np2 = 0.163
    Mauchly W = 0.884  p = 0.097   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a antibiotic x day mixed design we analyzed CRP for 40 patients (120 observations). The interaction is significant (F(2,76) = 7.42, p = 0.001, np2 = 0.163) ** -- the two groups' change across day differs in magnitude. The between-subjects main effect (regimen_A vs regimen_B) is F = 27.45, p < .001; the within-subjects main effect (day0/day3/day7) is F = 341.87, p < .001. Mauchly's test p = 0.097, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each day level. In infectious diseases, the mixed design is the standard analysis for comparing antibiotic regimens by the fall in inflammatory markers over days.

====================================================================================
