====================================================================================
MerQur - Fen_Matematik - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Fen_Matematik/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_mikrobiyoloji.xlsx
  >> SCENARIO (narration):
    We compiled the laboratory inventory of 280 samples. pH, temperature,
    concentration, absorbance and growth ratio were measured for each. Before any
    inferential test we want the overall picture of the measurements; so we begin
    with descriptive statistics.
  >> VARIABLE SELECTION:
    - Variables: pH
    - Variables: temperature_C
    - Variables: concentration_mM
    - Variables: absorbance
    - Variables: growth_ratio
    - Grouping (categorical): type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 samples
    pH: mean = 7.01   |   temperature: mean = 36.8 C   |   concentration: mean = 10.11 mM
    absorbance: mean = 0.587   |   growth ratio: mean = 0.453

>> COMMENTARY (narration):
    First we draw the overall picture of the sample set: 280 samples, average pH ~7.0, temperature ~37 C,
    concentration ~10 mM, absorbance 0.59 and growth ratio 0.45. This descriptive table lays the groundwork for every
    analysis that follows -- species/medium comparisons, physicochemical relationships, spatial sample pattern. In
    laboratory analytics, before any inferential test, summarizing the basic features of the measurements (pH,
    temperature, concentration, response) is essential both to audit data quality and to set experimental priorities.

====================================================================================

#2  Normality Tests
    file: 02_normality_chemical.xlsx
  >> SCENARIO (narration):
    We examine whether concentration, reaction rate and molecule weight are normally
    distributed. Because the validity of subsequent t-tests, ANOVA and correlation
    depends on this assumption, we test each variable separately.
  >> VARIABLE SELECTION:
    - Variables: concentration_mM
    - Variables: reaction_rate
    - Variables: molecule_weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    3 continuous variables tested:
    concentration_mM        : Shapiro-Wilk = 0.994  p = 0.542   KS p = 0.647   Normal
    reaction_rate           : Shapiro-Wilk = 0.900  p < .001    KS p = 0.001   Not Normal
    molecule_weight         : Shapiro-Wilk = 0.842  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three continuous variables -- concentration, reaction rate and molecule weight -- are normally
    distributed. The result splits instructively: concentration is normal (p = 0.54), but reaction rate and molecule
    weight deviate significantly from normality (p < .001). This is typical in chemical data -- reaction rates and
    weight distributions are often right-skewed. Practical upshot: we can safely use parametric tests (t-test, ANOVA,
    Pearson) on concentration; for reaction rate and molecule weight, nonparametric methods (Mann-Whitney/
    Kruskal-Wallis) or a transform are more appropriate. The normality check is a critical preliminary step that
    decides, per variable, which test family fits.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_pH.xlsx
  >> SCENARIO (narration):
    We investigate whether the samples' mean pH differs from the neutral reference
    value of 7.0. With one group and a fixed reference, the one-sample t-test is
    appropriate.
  >> VARIABLE SELECTION:
    - Test variable: pH
    - Test value (mu): 7.0

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = -1.863   p = 0.066 ns   Cohen d = -0.186 (Negligible)
    Mean pH = 6.96   (test mu = 7.0, neutral reference)   H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the samples' mean pH against the neutral reference value of 7.0. The result is not significant: mean
    6.96, statistically indistinguishable from 7.0 -- t(99) = -1.86, p = 0.066, negligible effect (d = -0.19). The
    samples can be considered neutral on average; the small observed deviation is explainable by chance. The
    one-sample t-test is the right way to compare a measurement against a known reference/standard (neutral pH, target
    concentration); as here, a "no difference" result is also valuable for calibration and quality control.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent-Samples t-Test
    file: 04_independent_t_catalyst.xlsx
  >> SCENARIO (narration):
    We compare the mean yield of two catalyst groups. With two separate groups and a
    continuous measure, the independent-samples t-test is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): catalyst
    - Test variable: yield_pct

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = 6.060   p < .001 ***   Cohen d = 1.131 (Large)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the mean yield (%) of two catalyst groups. The difference is significant and large: t(113) = 6.06,
    p < .001, d = 1.13. It shows the between-group difference is too pronounced to be chance and is also practically
    noteworthy. The independent-samples t-test is the standard way to compare the means of two separate groups (two
    catalysts, two methods, two conditions) on a continuous measure; it is a fundamental tool for detecting
    intervention/condition differences in the lab.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired-Samples t-Test
    file: 05_paired_t_cell_growth.xlsx
  >> SCENARIO (narration):
    We compare growth measured before and after an intervention in the same cells.
    Since the measures are paired, the paired t-test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: growth_before_um
    - 2nd measure / group: growth_post_um

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(49) = -18.19   p < .001 ***   Cohen d_z = -2.573 (Large)
    Mean diff (before - after) = -2.23 um   H0 REJECTED

>> COMMENTARY (narration):
    We paired and compared growth (um) measured before and after an intervention in the same cells. The result is
    very strong: mean difference -2.2 um, t(49) = -18.19, p < .001, d_z = -2.57, a huge effect. Growth rose
    significantly and substantially after the intervention. The paired t-test compares two timed measures on the same
    unit (before/after); by isolating individual change it is more powerful than the independent test and is the right
    way to measure an intervention's effect in the lab.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_type_reproduction.xlsx
  >> SCENARIO (narration):
    We compare the effect of four types on mean reproduction rate. With more than
    two groups, one-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: reproduction_rate
    - Factor (categorical): type

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,136) = 45.47   p < .001 ***   eta^2 = 0.501   H0 REJECTED

>> COMMENTARY (narration):
    We compared mean reproduction rate across four types. The result is significant and very strong: F(3,136) = 45.47,
    p < .001, eta^2 = 0.50 -- so half of reproduction-rate variance comes from type differences. At least one type
    differs significantly. One-way ANOVA compares the means of more than two groups at once (avoiding the error
    inflation of many t-tests); in the lab it is the core method for comparing different type/medium/condition
    performance. Which pairs differ is then determined by post-hoc tests.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova_type_medium.xlsx
  >> SCENARIO (narration):
    We examine the main effects and interaction of type and medium on yield
    simultaneously. With two categorical factors, two-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: yield
    - Factor (categorical): type
    - 2nd Factor: medium

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    type (main effect)   : F(2) = 219.69   p < .001 ***   eta^2p = 0.720
    medium (main effect) : F(2) = 28.16    p < .001 ***   eta^2p = 0.248

>> COMMENTARY (narration):
    We examined two factors at once: how do type and medium affect yield? Both main effects are very strong and
    significant (type: F(2) = 219.69, eta^2p = 0.72; medium: F(2) = 28.16, eta^2p = 0.25). The dominant driver of
    yield is type, but medium also makes an independent contribution. The power of two-way ANOVA is that it tests both
    factors and their interaction in a single model -- isolating each factor's pure effect with the other controlled.
    It is ideal for answering "which variable really makes a difference?" in the lab.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated-Measures ANOVA
    file: 08_repeated_anova_enzyme_temperature.xlsx
  >> SCENARIO (narration):
    We compare an enzyme measure over four consecutive conditions in the same units.
    With repeated measures on the same unit, repeated-measures ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: measurement_1
    - Repeated measures: measurement_2
    - Repeated measures: measurement_3
    - Repeated measures: measurement_4

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,177) = 108.40   p < .001 ***   eta^2p = 0.648   n = 60   H0 REJECTED

>> COMMENTARY (narration):
    We compared an enzyme measure across four consecutive conditions (e.g. rising temperature) in the same 60 units.
    The result is very strong: F(3,177) = 108.40, p < .001, eta^2p = 0.65 -- the between-condition difference is huge
    and most of the effect is condition-related. The measure changes significantly across conditions.
    Repeated-measures ANOVA compares three or more measurements on the same unit; by holding individual differences
    constant it yields high statistical power and is the right choice for multi-condition/multi-time measurement
    tracking (temperature series, time points) in the lab.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_type_3DV.xlsx
  >> SCENARIO (narration):
    We test type's effect on molecule weight, pH and absorbance at once. With
    several correlated dependent variables, MANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: molecule_weight
    - Dependent variable: pH
    - Dependent variable: absorbance
    - Factor (categorical): type

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.237   F(6,230) = 40.39   p < .001 ***   n = 120
    Dependent: molecule_weight, pH, absorbance   |   Factor: type   H0 REJECTED

>> COMMENTARY (narration):
    We tested type's effect on three dependent variables (molecule weight, pH, absorbance) simultaneously. With Wilks'
    Lambda = 0.24, F(6,230) = 40.39, p < .001, the effect is very strong. MANOVA examines several correlated outcomes
    in one test, both preventing the error inflation of many separate ANOVAs and capturing the joint information the
    variables carry together. In the lab, when type affects not a single indicator but the physicochemical "bundle"
    together, MANOVA reveals this multivariate difference as a single decision.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova_dose_activity.xlsx
  >> SCENARIO (narration):
    We compare post-intervention activity across groups while controlling baseline
    activity as a covariate. With a confounding continuous variable, ANCOVA is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: activity_after
    - Factor (categorical): group
    - Covariate: activity_baseline

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group effect significant   eta^2p = 0.867 (very large)   covariate: activity_baseline   H0 REJECTED

>> COMMENTARY (narration):
    We compared post-intervention activity (activity_after) across groups while controlling baseline activity
    (activity_baseline) as a covariate. This rules out the objection that "the groups differed at baseline" and
    measures the pure group effect: the effect is very large (eta^2p = 0.87). ANCOVA makes the group comparison fair
    by statistically holding a confounding continuous variable constant -- it is the answer to "once we equalize the
    starting level, does the dose/intervention still make a difference?" and is the standard tool in
    baseline/post-measure designs in the lab.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci_half_life.xlsx
  >> SCENARIO (narration):
    For mean isotope half-life we build a confidence interval via resampling, with
    no distributional assumption. For skewed data, bootstrap is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: half_life_minute

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean (half-life, min) = 37.64
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    For the mean isotope half-life (min) we produced a 95% confidence interval via resampling -- with no
    distributional assumption: observed mean 37.64 min. Bootstrap builds the sampling distribution of the statistic
    empirically by resampling the data thousands of times from itself; it is a reliable way to give a confidence
    interval when normality does not hold or no formula is known. In the lab it provides more robust estimates than
    classic t-intervals for skewed quantities (half-life, waiting time, damage).

====================================================================================

#12  Permutation Test
    file: 12_permutation_method_yield.xlsx
  >> SCENARIO (narration):
    We test the yield difference between two groups via permutation, with no
    distributional assumption. For small samples/odd distributions, permutation is
    appropriate.
  >> VARIABLE SELECTION:
    - Test variable: yield_pct
    - Grouping (categorical): group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = -4.897   p < .001 ***   significant mean difference between the two groups   H0 REJECTED

>> COMMENTARY (narration):
    We tested the yield (%) difference between two groups without any distributional assumption, via permutation: the
    null distribution was built by randomly swapping group labels, and the observed difference turned out very rare in
    that distribution (p < .001). Because the permutation test is exact and distribution-free, it is a safe
    alternative to the parametric t-test for small samples or odd distributions; in the lab it is a robust choice for
    group comparisons where assumptions are in doubt.

====================================================================================

#13  Multiple Comparison
    file: 13_multiple_comparison_catalyst.xlsx
  >> SCENARIO (narration):
    We take the p-values of eight catalyst comparisons together and apply
    multiple-testing correction. With many tests, p-adjustment is appropriate.
  >> VARIABLE SELECTION:
    - Variables: catalyst
    - Variables: conversion

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Raw: 6/8 significant
    After Bonferroni / Holm / BH (FDR): 5/8 significant

>> COMMENTARY (narration):
    We took the p-values of eight separate catalyst comparisons together and applied multiple-testing correction. Raw,
    6 comparisons were significant; after Bonferroni/Holm/BH that dropped to 5. When many tests are run, the rate of
    false positives that look "significant" by chance alone inflates; correction methods tighten the threshold to
    control this error. In the lab, when many conditions/catalysts are compared at once, some "significant" conclusions
    reached without correction can be misleading -- this step protects inference reliability.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney_habitat_richness.xlsx
  >> SCENARIO (narration):
    We compare the type-richness distribution of two habitats. Since normality
    fails, Mann-Whitney is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): habitat
    - Test variable: type_richness

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    U = 1062.00   p = 0.149 ns   r = -0.18   H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the type-richness distribution of two independent habitats based on ranks rather than means: U = 1062,
    p = 0.149, small effect (r = -0.18) -- no significant difference. Mann-Whitney is the nonparametric counterpart of
    the t-test; when normality fails or the scale is count/ordinal, it is the right way to compare two groups. Here the
    "no difference" result tells us the two habitats are similar in type richness.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank
    file: 15_wilcoxon_quality.xlsx
  >> SCENARIO (narration):
    We compare quality scores measured before/after in the same units, without
    assuming normality. For paired non-normal data, Wilcoxon is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: quality_before
    - 2nd measure / group: quality_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    W = 0.00   p < .001 ***   r = 0.89   n (non-zero) = 31   H0 REJECTED

>> COMMENTARY (narration):
    We compared quality scores measured before/after in the same units -- without assuming normality, on a rank basis:
    W = 0, p < .001, very large effect (r = 0.89). Quality changed consistently and strongly after the intervention.
    Wilcoxon is the nonparametric counterpart of the paired t-test; it is the right choice for ordinal or non-normal
    before/after measures. In the lab it gives reliable results for pre/post-improvement quality measurements.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_pH_activity.xlsx
  >> SCENARIO (narration):
    We compare the enzyme-activity distribution across three pH levels. Since
    normality/variance homogeneity fails, Kruskal-Wallis is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): ph_level
    - Test variable: enzyme_activity

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 18.80   p < .001 ***   eta^2_H = 0.40   H0 REJECTED

>> COMMENTARY (narration):
    We compared the enzyme-activity distribution across three pH levels by ranks rather than means: H(2) = 18.80,
    p < .001, eta^2_H = 0.40, a strong effect. At least one group differs significantly. Kruskal-Wallis is the
    nonparametric counterpart of one-way ANOVA; when normality or variance homogeneity fails, it is the right way to do
    multi-group comparison. In the lab it is robust for detecting group differences on skewed or ordinal measures.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman_method_quality.xlsx
  >> SCENARIO (narration):
    We compare quality scores measured under four conditions in the same units. As a
    nonparametric repeated measure, Friedman is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: condition_A
    - Repeated measures: condition_B
    - Repeated measures: condition_C
    - Repeated measures: condition_D

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 39.08   p < .001 ***   Kendall W = 0.434   n = 30   H0 REJECTED

>> COMMENTARY (narration):
    We compared quality scores measured under four conditions (condition_A..D) in the same 30 units -- as a
    nonparametric repeated measure: chi^2(3) = 39.08, p < .001, Kendall W = 0.43 (moderate concordance). The
    between-condition difference is significant. Friedman is the nonparametric counterpart of repeated-measures ANOVA;
    it is the right choice for ordinal or non-normal paired multi-condition measures. In the lab it is used to compare
    the same sample's rankings across different methods/conditions.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_experiment_achievement.xlsx
  >> SCENARIO (narration):
    We test the proportion of successful experiments against an expected 50%. With a
    binary outcome and a theoretical proportion, the binomial test is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: successful
    - Expected proportion: 0.50
    - Success value: 1

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed proportion = 0.736   (expected p = 0.50)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested the proportion of successful experiments against an expected 50%: observed proportion 73.6%, well above
    expectation (p < .001). The success rate is too high to be chance. The binomial test is the exact method for
    comparing the observed proportion of a binary (yes/no) outcome with a theoretical proportion; in the lab it
    directly tests whether experiment success/positive-result rates meet a target or a 50:50 expectation.

====================================================================================

#19  Sign Test
    file: 19_sign_test_viscosity.xlsx
  >> SCENARIO (narration):
    We look at the direction of before/after viscosity in the same units. When only
    directional information is reliable, the sign test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: viscosity_before
    - 2nd measure / group: viscosity_post

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive diffs (before > after) = 27   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We looked at the direction of change in viscosity measured before/after in the same units: the large majority of
    changes go in one direction, p < .001 for a significant directional change. The sign test uses only the direction
    of the difference (increase/decrease), not its magnitude, so it is the before/after test that requires the fewest
    assumptions. In the lab it is a robust choice when the measure is skewed or only directional information is
    reliable (did viscosity go up or down).

====================================================================================

#20  Runs Test
    file: 20_runs_test_measurement.xlsx
  >> SCENARIO (narration):
    We test whether the measurement sequence (around the median) is random. For
    sequence randomness, the runs test is appropriate.
  >> VARIABLE SELECTION:
    - Column: measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 28   Z = -0.781   p = 0.435 ns   data may be considered random

>> COMMENTARY (narration):
    We tested whether the binary sequence of measurements (around the median) is random: 28 runs, Z = -0.78, p = 0.44
    -- no pattern, the sequence is random. The runs test checks whether values in a sequence form a systematic pattern
    (clusters, cycles, trend). In the lab it is used to determine whether systematic patterns or randomness dominate in
    measurement/quality/production series, and in checking process control and the independence assumption.

====================================================================================

#21  Chi-Square Independence
    file: 21_chisquare_type_color.xlsx
  >> SCENARIO (narration):
    We test whether type and color are related. With two categorical variables,
    chi-square is appropriate.
  >> VARIABLE SELECTION:
    - Variables: type
    - Variables: color

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(2) = 43.20   p < .001 ***   Cramer's V = 0.40 (Strong)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether two categorical variables -- type and color -- are related: chi^2(2) = 43.20, p < .001, Cramer's
    V = 0.40, a strong dependency. The categories are not independent; they vary together. The chi-square test of
    independence detects the relationship between two qualitative variables from a cross-tab; in the lab it is the core
    method for revealing categorical relationships such as type-property, condition-outcome. Cramer's V measures the
    practical strength of the relationship.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_category.xlsx
  >> SCENARIO (narration):
    We test whether a four-category variable's observed distribution fits an equal
    expected distribution. For one categorical variable, goodness-of-fit is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: category

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 58.68   p < .001 ***   N = 200, k = 4 (expected: equal distribution)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the observed distribution of a four-category variable (category) fits an equal (1/k) expected
    distribution: chi^2(3) = 58.68, p < .001 -- the categories are not equally distributed, some are clearly more
    frequent. The goodness-of-fit test compares the observed frequencies of a single categorical variable with a
    theoretical expectation (equal proportions, a known ratio); in the lab it tests whether class/category
    distributions match an expected profile.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_method.xlsx
  >> SCENARIO (narration):
    In a 2x2 table we test the method-achievement relationship exactly, due to small
    cell frequencies. For few observations, Fisher is appropriate.
  >> VARIABLE SELECTION:
    - Variables: method
    - Variables: achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher exact p = 0.655 ns   Odds Ratio = 0.50   Phi = 0.137   H0 NOT REJECTED

>> COMMENTARY (narration):
    In a 2x2 cross-tab (method x achievement) we tested the relationship exactly -- using Fisher instead of chi-square
    because of small cell frequencies: p = 0.655, no significant relationship (Phi = 0.14). Fisher's exact test is the
    right choice when expected frequencies are low and the chi-square approximation is unreliable; it computes the
    probability exactly rather than approximately. Here the "no relationship" result tells us method is independent of
    achievement.

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_test_comparison.xlsx
  >> SCENARIO (narration):
    We compare the binary outcome of two methods paired in the same units. For
    paired binary change, McNemar is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: test_a
    - 2nd measure / group: test_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar (exact) min(b,c) = 5   p = 0.096 ns   discordant b = 13, c = 5   OR = 2.60   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested the direction of change in a binary outcome of two methods (test_a vs test_b) paired in the same units:
    although the changes are imbalanced (b = 13, c = 5), p = 0.096 is non-significant. McNemar tests the before/after
    (or two-method) binary status change in the same unit and looks only at discordant pairs. In the lab it is the
    right tool for detecting whether two methods' binary outcomes differ systematically; the exact (binomial) version
    is used for small discordant counts.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_observer.xlsx
  >> SCENARIO (narration):
    We measure the agreement of two observers classifying the same samples. For
    categorical agreement, kappa is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: observer_a
    - 2nd measure / group: observer_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Kappa = 0.760 (substantial agreement)

>> COMMENTARY (narration):
    We measured the agreement of two observers (observer_a vs observer_b) classifying the same samples into the same
    categories -- excluding chance agreement: kappa = 0.76, substantial. Unlike raw percent agreement, kappa reports
    categorical agreement after removing the chance-agreement share, so it is more honest. In the lab it is the
    standard index for measuring how consistent two observers' classifications (species ID, quality class) are.

====================================================================================

#26  Cochran-Mantel-Haenszel
    file: 26_cmh_medium_group_death.xlsx
  >> SCENARIO (narration):
    We test the group-death relationship controlling for medium strata. For a
    stratum-controlled relationship, CMH is appropriate.
  >> VARIABLE SELECTION:
    - Variables: group
    - Variables: death
    - Stratum: medium

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi^2(1) = 25.63   p < .001 ***   Common OR (MH) = 3.01 [1.95, 4.62]
    Breslow-Day p = 0.520 (OR homogeneous)   H0 REJECTED

>> COMMENTARY (narration):
    We tested the group-death relationship while controlling for medium strata: the common Odds Ratio across strata =
    3.01 [1.95, 4.62], p < .001; Breslow-Day p = 0.52 means this relationship is consistent across all strata. CMH
    measures the pure strength of the association by holding a confounding stratum variable constant (preventing
    Simpson's paradox). In the lab it is ideal for robustly estimating a group-outcome relationship while controlling
    medium/batch differences.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_3yollu.xlsx
  >> SCENARIO (narration):
    We model the joint relationship structure of three categorical variables. For
    more than two categorical dimensions, log-linear is appropriate.
  >> VARIABLE SELECTION:
    - Variables: F1
    - Variables: F2
    - Variables: F3

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 71.97   Pearson chi^2 = 0.506   joint relationship of 3 categorical variables (F1, F2, F3) modeled

>> COMMENTARY (narration):
    We examined the joint relationship structure of three categorical variables (F1, F2, F3) with a log-linear model.
    The model explains cell frequencies via main effects and interactions; it reveals which pairs/triples of variables
    vary together. It is the multivariable generalization of the two-way cross-tab. In the lab it is used to analyze
    the joint dependency pattern of more than three categorical dimensions (type x condition x outcome), balancing
    parsimony with AIC to select the most explanatory structure.

====================================================================================

#28  Cross-Tabulation
    file: 28_cross_age_category.xlsx
  >> SCENARIO (narration):
    We examine age group and category in a cross-tab. For the relationship of two
    qualitative variables, cross-tab is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age_group
    - Variables: category

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(6) = 19.58   p = 0.003 **   Cramer's V = 0.167 (Moderate)   H0 REJECTED

>> COMMENTARY (narration):
    We examined two categorical variables (age group x category) in a cross-tab: chi^2(6) = 19.58, p = 0.003, V = 0.17
    -- a significant but moderate-to-weak relationship. Cross-tab + chi-square shows the direction and strength of the
    association between two qualitative variables at the cell level. In the lab it is used to describe class-category
    relationships and to correctly interpret significant but practically modest (V = 0.17) associations.

====================================================================================

#29  Multiple Response - Frequency
    file: 29_mr_frequency_elements.xlsx
  >> SCENARIO (narration):
    We analyze a multi-select elements question. For a multi-select question,
    multiple-response frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: elements (coklu yanit / multi-response)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)   multi-select "elements" question

>> COMMENTARY (narration):
    We analyzed a question where several options could be checked (elements present in a sample): each of 250 samples
    listed one or more elements. Multiple-response frequency analysis gives the count of checks per option and both the
    response and case percentages separately (percentages sum to over 100, because one sample contains several). In the
    lab it is the standard method for correctly summarizing multi-select content questions (components present, species
    detected).

====================================================================================

#30  Multiple Response - Cross-Tab
    file: 30_mr_categorical_type_element.xlsx
  >> SCENARIO (narration):
    We cross-tabulate the multi-response elements question by type. To break a
    multi-select by a category, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: elements (coklu yanit / multi-response)
    - Grouping (categorical): type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: type   |   multi-response "elements" x type

>> COMMENTARY (narration):
    We cross-tabulated the multi-response elements question by a single category (type): we compared the per-element
    check rates for each type. A multiple-response cross-tab answers "does one group contain certain options more often
    than another?". In the lab it is used to compare the multi-component profiles of types/groups (elements they
    contain, components detected).

====================================================================================

#31  Multiple Response x Multiple Response
    file: 31_mr_mr_two_group.xlsx
  >> SCENARIO (narration):
    We cross-tabulate two multi-response questions against each other. For
    many-to-many co-occurrence, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: group_1 (coklu / multi)
    - Variables: group_2 (coklu / multi)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180   two multi-responses (group_1 x group_2) cross-tabulated

>> COMMENTARY (narration):
    We cross-tabulated two separate multi-response questions against each other. This is the most complex form of
    tabulation: on both axes a unit contributes to more than one cell. Multiple-response by multiple-response reveals
    many-to-many co-occurrences such as "which options appear together with which options?". In the lab it is used to
    examine the matching pattern of nested component bundles (content x property).

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q_condition.xlsx
  >> SCENARIO (narration):
    We test whether a binary outcome varies across conditions in the same units. For
    3+ repeated binary measures, Cochran's Q is appropriate.
  >> VARIABLE SELECTION:
    - Columns: kosul-bazli ikili sutunlar / condition-wise binary columns

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000 ns   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the binary outcome (positive response yes/no) measured for different conditions in the same units
    varies significantly from condition to condition: Q = 0, p = 1.00 -- no difference across conditions. Cochran's Q
    is the generalization of McNemar to more than two repeated conditions; it compares 3+ binary measures in the same
    unit. In the lab it is the right method for comparing whether the same samples give a positive result across
    multiple conditions/methods.

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5fizikokimyasal.xlsx
  >> SCENARIO (narration):
    We examine all pairwise correlations among five physicochemical variables. For
    relationship direction and strength, the correlation matrix is appropriate.
  >> VARIABLE SELECTION:
    - Variables: pH
    - Variables: temperature_C
    - Variables: concentration_mM
    - Variables: activity
    - Variables: yield_pct

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables: pH, temperature_C, concentration_mM, activity, yield_pct
    Strong relationship (|r| >= 0.75): 1 pair -> pH with temperature_C: r = +0.789 (p < .001)

>> COMMENTARY (narration):
    We computed all pairwise Pearson correlations among five physicochemical variables: only pH and temperature were
    strongly positively related (r = 0.79, p < .001); the other pairs were weaker. So pH and temperature rise together,
    while the remaining measures carry largely separate information. The correlation matrix summarizes the direction
    and strength of relationships at a glance; in the lab it is the first step in spotting overlap among indicators
    (multicollinearity risk) and seeing which variables truly give separate signals.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Agreement
    file: 34_bland_altman_device.xlsx
  >> SCENARIO (narration):
    We examine the agreement of two devices measuring the same quantity. For device
    interchangeability, Bland-Altman is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: device_A_measurement
    - 2nd measure / group: device_B_measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Device A (device_A): mean = 51.85   |   Device B (device_B): mean = 50.33
    Bias (mean difference) ~ 1.52   |   limits of agreement (LoA) computed

>> COMMENTARY (narration):
    We examined how well two devices measuring the same quantity (device_A vs device_B) agree: the mean systematic
    difference (bias) is ~1.5 units, and 95% limits of agreement were reported. Unlike correlation, Bland-Altman
    answers "can the two devices be used interchangeably?" -- high correlation does not mean agreement, there may be a
    systematic shift. In the lab it is the standard method for testing the interchangeability of two measuring
    instruments/methods.

====================================================================================

#35  Effect Size
    file: 35_effect_size_method.xlsx
  >> SCENARIO (narration):
    We measure the practical size of the yield difference between two methods. For
    importance beyond the p-value, effect size is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: yield
    - Grouping (categorical): method

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.536 (medium effect)   yield difference between two methods (method)

>> COMMENTARY (narration):
    We measured the practical size of the yield difference between two methods, independent of the p-value, with
    Cohen's d: d = -0.54, a medium effect. The p-value answers "is there a difference?"; effect size answers "how
    important is it?". Because large samples can produce significant but trivial differences, reporting effect size is
    essential. In the lab this measure clarifies whether the difference between two methods/protocols is practically
    noteworthy.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_physical_chemistry.xlsx
  >> SCENARIO (narration):
    We examine the joint structure between the physical-properties set and the
    chemical-properties set. For the relationship between two multivariate sets, CCA
    is appropriate.
  >> VARIABLE SELECTION:
    - X variables: density
    - X variables: viscosity
    - X variables: conductivity
    - X variables: color
    - X variables: stability
    - Y variables: ph
    - Y variables: purity
    - Y variables: yield
    - Y variables: activity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.976   chi^2(20) = 1233.62   p < .001 ***
    CC2: r = 0.958   chi^2(12) = 795.90    p < .001 ***
    Set-X (physical): density, viscosity, conductivity, color, stability  |  Set-Y (chemical): ph, purity, yield, activity

>> COMMENTARY (narration):
    We resolved the joint structure between two multivariate measure sets -- physical properties (density, viscosity,
    conductivity, color, stability) and chemical properties (pH, purity, yield, activity) -- with canonical
    correlation. The first two canonical functions are very strong (r = 0.976 and 0.958, p < .001): the two sets are
    intensely related. CCA answers "how is one variable set related to another?" in a single step -- it is the
    multivariate-on-both-sides version of multiple regression. In the lab it reveals the latent relationship structure
    between the physical bundle and the chemical bundle.

====================================================================================

#37  Correspondence Analysis
    file: 37_ca_acidity_appearance.xlsx
  >> SCENARIO (narration):
    We map the relationship between acidity and appearance. To see the structure of
    two categorical variables, CA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: acidity
    - Variables: appearance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.109   (2 dimensions)   acidity x appearance

>> COMMENTARY (narration):
    We mapped the relationship in the cross-tab of two categorical variables (acidity x appearance) into a visual space
    with correspondence analysis: total inertia 0.109 (relatively weak relationship) resolved over two dimensions. CA
    positions the chi-square relationship in a two-dimensional space, showing which categories are close (co-occurring).
    In the lab it is powerful for visually interpreting the structure of qualitative relationships such as
    acidity-appearance, type-property.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_18degisken.xlsx
  >> SCENARIO (narration):
    We cluster eighteen items (physical/chemical/biological) by their similarity. To
    find the latent dimension structure, VarClus is appropriate.
  >> VARIABLE SELECTION:
    - Variables: physical_m1..6, chemistry_m1..6, bio_m1..6 (18 madde / items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: physical_m1..6, chemistry_m1..6, bio_m1..6)

>> COMMENTARY (narration):
    We clustered eighteen measurement items (physical, chemical, biological, 6 each) by how related they are: they
    grouped into 3 main dimensions -- most likely matching the natural physical/chemical/biological groups. VarClus
    groups the VARIABLES, not the observations -- by placing highly correlated items in the same cluster it reveals the
    latent dimensional structure of the data set. In the lab it is practical for reducing a long indicator set to a few
    core dimensions and for spotting redundant items.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_yield.xlsx
  >> SCENARIO (narration):
    We model yield with four predictors. To explain a continuous outcome with
    multiple variables, multiple regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: yield_pct
    - Predictor(s): pH
    - Predictor(s): temperature_C
    - Predictor(s): concentration_mM
    - Predictor(s): catalyst_g_L

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.773   Adj. R^2 = 0.768   predictors: pH, temperature_C, concentration_mM, catalyst_g_L

>> COMMENTARY (narration):
    We modeled yield (yield_pct) with four predictors (pH, temperature, concentration, catalyst) at once: the model
    explains 77.3% of variance (Adj. R^2 = 0.77) -- very strong explanatory power. Multiple regression gives each
    predictor's pure contribution to yield with the others held constant; thus it answers "which factor really raises
    yield?" while controlling confounders. In the lab it is the core method for identifying the drivers of yield and
    for prediction.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_reaction_achievement.xlsx
  >> SCENARIO (narration):
    We model reaction success (successful/failed) with three predictors. For a
    binary outcome, logistic regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: reaction_successful
    - Predictor(s): temperature_C
    - Predictor(s): catalyst
    - Predictor(s): pH_deviation

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 = 0.202   binary outcome: reaction_successful   predictors: temperature_C, catalyst, pH_deviation

>> COMMENTARY (narration):
    We modeled a binary outcome (reaction successful/failed) with three predictors: the model has moderate explanatory
    power (pseudo R^2 = 0.20) and gives each predictor's effect on the odds. Logistic regression replaces linear
    regression when the outcome is binary; coefficients are converted to Odds Ratios to read "how many times does the
    success odds change per unit increase in this variable?". In the lab it is the core model for predicting yes/no
    outcomes such as reaction/experiment success.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Count (Poisson) Regression
    file: 41_poisson_type_count.xlsx
  >> SCENARIO (narration):
    We model the type count with temperature and moisture. For a count outcome,
    Poisson regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: type_count
    - Predictor(s): temperature_C
    - Predictor(s): moisture_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 849.61   Deviance = 175.03   count outcome: type_count   predictors: temperature_C, moisture_pct

>> COMMENTARY (narration):
    We modeled a count variable (type count) with temperature and moisture. Poisson regression is the right model when
    the outcome is a count (0,1,2,... items); linear regression is unsuitable because it can produce negative/fractional
    predictions. Coefficients give the effect on the count rate. In the lab/ecology it is used to explain count outcomes
    such as species count, colony count, event frequency with environmental conditions.

====================================================================================

#42  Multinomial Logistic
    file: 42_multinomial_crystal.xlsx
  >> SCENARIO (narration):
    We model the multi-category crystal structure with two predictors. For a nominal
    multi-class outcome, multinomial logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: crystal_structure
    - Predictor(s): pH
    - Predictor(s): temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 367.86   Reference class: 'structure-A'   outcome: crystal_structure (>2 categories)   predictors: pH, temperature_C

>> COMMENTARY (narration):
    We modeled a more-than-two-category outcome (crystal structure) with two predictors (pH, temperature). Multinomial
    logistic compares each category against a reference class (here 'structure-A') with a separate logistic equation;
    coefficients are read as "as X increases, how does the chance of being in this structure change relative to the
    reference?". When the outcome is nominal with more than two classes (crystal type, phase, category) it is the right
    choice. In the lab it is the standard model for multi-option class prediction.

====================================================================================

#43  Ordinal Logistic
    file: 43_ordinal_quality.xlsx
  >> SCENARIO (narration):
    We model the ordinal quality level with energy and purity. For an ordinal
    outcome, ordinal logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: quality_level
    - Predictor(s): energy_kJ
    - Predictor(s): purity_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 344.52   ordinal outcome: quality_level   predictors: energy_kJ, purity_pct

>> COMMENTARY (narration):
    We modeled an ordinal outcome (quality level: low/medium/high) with energy and purity. Ordinal logistic uses the
    ORDER information between categories (which multinomial ignores); with a "proportional odds" assumption it explains
    all thresholds with one coefficient set. In the lab it is the right and more powerful choice for modeling naturally
    ordered outcomes such as quality level, grade, class.

====================================================================================

#44  PLS Regression
    file: 44_pls_sensor_concentration.xlsx
  >> SCENARIO (narration):
    We predict concentration from 12 sensor readings. For many highly correlated
    predictors, PLS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: concentration_mM
    - Predictor(s): sensor_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 (train) = 0.865   R^2 (5-fold CV) = 0.848   outcome: concentration_mM   predictors: sensor_01..12 (12 sensors)

>> COMMENTARY (narration):
    We predicted concentration from 12 sensor readings. Because the sensor readings are highly correlated
    (multicollinearity), classic regression becomes unstable; PLS reduces them to a few latent components and regresses
    on those. With cross-validated R^2 = 0.85, the model is both strong and generalizable. PLS is ideal when predictors
    are numerous or highly correlated; in chemometrics/spectroscopy (estimating concentration from multi-sensor
    readings) it is widely used.

====================================================================================

#45  Probit Regression
    file: 45_probit_dose_response.xlsx
  >> SCENARIO (narration):
    We estimate the probability of death from dose with a probit model. For a binary
    dose-response outcome, probit is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: death
    - Predictor(s): dose_mg_L

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 120.01   Pseudo R^2 (McFadden) = 0.433   binary outcome: death   predictor: dose_mg_L

>> COMMENTARY (narration):
    We estimated the dose-response relationship -- the probability of death from dose (dose_mg_L) -- with a probit
    model: the model is strong (pseudo R^2 = 0.43). Probit, like logistic, applies to binary outcomes; the difference
    is that its link function is the normal distribution. Probit is classic in dose-response studies (LD50 estimation
    in toxicology). In the lab it is a robust choice for modeling the relationship of a binary response with a
    continuous stimulus (dose, concentration).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_production.xlsx
  >> SCENARIO (narration):
    We model a threshold-piled production variable with purity and price. For a
    censored outcome, tobit is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: production_kg
    - Predictor(s): purity_pct
    - Predictor(s): price_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 2354.12   Left censor   outcome: production_kg (censored)   predictors: purity_pct, price_TL

>> COMMENTARY (narration):
    We modeled a lower-bounded (threshold-piled) production variable with purity and price. Tobit is for "censored"
    dependent variables that pile up at a threshold; ordinary regression gives biased estimates by ignoring this
    pile-up. In the lab/production it is the right model for floor-bounded outcomes (zero production, below-threshold
    measurement); coefficients reflect the true (uncensored) relationship.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_reagent_yield.xlsx
  >> SCENARIO (narration):
    We model yield with two reagent concentrations in a Bayesian framework. To
    express uncertainty probabilistically, Bayesian regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: yield
    - Predictor(s): reagent_A_mM
    - Predictor(s): reagent_B_mM

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2 posterior mean = 9.39 (sd = 1.34)   outcome: yield   predictors: reagent_A_mM, reagent_B_mM
    coefficient posteriors + P(beta>0) reported

>> COMMENTARY (narration):
    We modeled yield with two reagent concentrations in a Bayesian framework: instead of point estimates we obtained
    each coefficient's full posterior distribution and the "probability the effect is positive". The Bayesian approach
    expresses uncertainty directly in probability language and can incorporate prior knowledge. In the lab, when the
    sample is small or prior-experiment information is valuable, it offers intuitive interpretations like "the effect
    is probably positive".

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear_kinetik.xlsx
  >> SCENARIO (narration):
    We model the S-shaped kinetics of product concentration over time via a logistic
    growth curve. For a curved relationship, nonlinear regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: product_concentration_mM
    - Predictor(s): time_min
    - Value: fonksiyon/function: logistic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Function: y = K / (1 + exp(-r*(x-x0)))  (logistic growth)   R^2 = 0.978

>> COMMENTARY (narration):
    We modeled the kinetic relationship between product concentration (product_concentration_mM) and time (time_min)
    not as a straight line but as an S-shaped logistic growth curve: the fit is very high (R^2 = 0.98). Nonlinear
    regression fits a theoretical function form directly to the data when the relationship is curved (saturation,
    threshold, exponential growth) and makes the parameters (ceiling K, rate r, inflection x0) interpretable. In the
    lab it is the right tool for modeling reaction kinetics, growth curves and saturation processes.

====================================================================================

#49  Ridge Regression
    file: 49_ridge_spectrum.xlsx
  >> SCENARIO (narration):
    We predict concentration from 15 spectral points with ridge. For
    multicollinearity, ridge is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: concentration
    - Predictor(s): spectrum_p01..15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 1.0   R^2 = 0.789   outcome: concentration   predictors: spectrum_p01..15 (15 spectral points)

>> COMMENTARY (narration):
    We predicted concentration from 15 spectral points with ridge regression (R^2 = 0.79). Ridge adds an L2 penalty to
    shrink all coefficients in magnitude without zeroing them; this prevents the instability caused by high correlation
    among neighboring spectral points (multicollinearity). Because spectral data are inherently highly correlated, ridge
    is common in chemometrics. In the lab it is preferred for prediction with many co-varying measurements.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso_30pred_5aktif.xlsx
  >> SCENARIO (narration):
    We model biological activity from 30 candidate features with lasso, selecting
    the important ones. For automatic variable selection, lasso is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: biological_activity
    - Predictor(s): x01..30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.1   R^2 = 0.715   16 of 30 features shrunk to zero (automatic variable selection)

>> COMMENTARY (narration):
    We modeled biological activity from 30 candidate features with lasso regression: R^2 = 0.72, and lasso shrank 16 of
    the 30 feature coefficients exactly to zero, selecting only the effective variables. This is its difference from
    ridge: because lasso can zero coefficients, it performs prediction and variable selection at the same time. In the
    lab it is very useful for automatically winnowing the "few truly important factors" out of many candidate
    indicators/features (a sparse, interpretable model).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_chemical.xlsx
  >> SCENARIO (narration):
    We test the catalyst -> intermediate -> final product chain. To resolve the
    intermediate mechanism, mediation analysis is appropriate.
  >> VARIABLE SELECTION:
    - Predictor(s): catalyst
    - Value: M (araci/mediator): intermediate_product_mM
    - Dependent variable: last_product_mM

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 1.093   95% CI [0.832, 1.346]   (excludes zero -> significant)
    X: catalyst  M: intermediate_product_mM  Y: last_product_mM

>> COMMENTARY (narration):
    We tested the chain "catalyst (X) -> intermediate product (M) -> final product (Y)": the indirect effect is 1.09,
    its 95% CI excludes zero -- so the catalyst's effect on the final product occurs substantially through the
    intermediate. Mediation analysis resolves "why/how does X affect Y?" through an intermediate mechanism. In the lab
    it is powerful for understanding which intermediate product/stage a reaction step's effect flows through.

====================================================================================

#52  Path Analysis
    file: 52_path_chemical_zincir.xlsx
  >> SCENARIO (narration):
    We test direct/indirect relationships as a single causal diagram. For a
    relationship network, path analysis is appropriate.
  >> VARIABLE SELECTION:
    - Value: yield ~ temperature_C + concentration_mM + activity
    - Value: activity ~ temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.917   RMSEA = 0.216   model: yield ~ temperature + concentration + activity; activity ~ temperature

>> COMMENTARY (narration):
    We tested the direct and indirect relationships among several variables as a single causal diagram. The fit indices
    (CFI = 0.92, RMSEA = 0.22) show the model fits the data partially, with room for improvement. Path analysis
    estimates the whole relationship network at once instead of separate regressions; it shows how variables affect
    each other and a common outcome. In the lab it is used to test theory-based relationship chains (temperature ->
    activity -> yield).

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_sample_time.xlsx
  >> SCENARIO (narration):
    We model repeatedly measured absorbance taking sample as a random effect. For
    repeated/nested data, LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: absorbance
    - Predictor(s): time_hour
    - Cluster: sample_id (random)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC = 0.761   Group variance (random intercept) = 0.0084   outcome: absorbance   fixed: time_hour   group: sample_id

>> COMMENTARY (narration):
    We modeled absorbance measured repeatedly in the same samples, taking sample identity as a random effect. ICC =
    0.76 is high: most of the absorbance variability comes from between-sample differences, while within-sample repeats
    are similar. LMM correctly handles dependency in nested/repeated (measurements within sample) data; it solves the
    "independence" assumption that ordinary regression violates via random effects. In the lab it is the right choice
    for repeated/time-series measurement data.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation.xlsx
  >> SCENARIO (narration):
    Instead of deleting missing data we fill it with 5 plausible value sets. To
    handle missingness without bias, multiple imputation is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: yield
    - Predictor(s): pH
    - Predictor(s): temperature_C
    - Predictor(s): concentration

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (imputation) = 5   outcome: yield   predictors: pH, temperature_C, concentration

>> COMMENTARY (narration):
    Instead of deleting missing data, we analyzed by producing 5 plausible value sets (m = 5) and combining the
    results. Multiple imputation -- unlike filling gaps with a single estimate (which ignores uncertainty) -- accounts
    for imputation uncertainty too, yielding unbiased estimates and correct standard errors. In the lab it is the
    modern standard for handling measurement dropouts/missing readings without shrinking the sample or distorting
    results.

====================================================================================

#55  GEE
    file: 55_gee_measurement.xlsx
  >> SCENARIO (narration):
    We model the measurement in repeated measures of the same samples. For a
    population-average effect, GEE is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: measurement
    - Predictor(s): visit
    - Cluster: sample_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coef = 0.052   p < .001 ***   QIC = 307.67   outcome: measurement   group: sample_id

>> COMMENTARY (narration):
    We modeled the measurement in repeated visit/time measures of the same samples; the visit effect is significant
    (b = 0.052, p < .001). GEE estimates the POPULATION-AVERAGE effect rather than individual effects in
    repeated/clustered data and corrects within-group correlation with a "working correlation structure". While LMM
    focuses on individual random effects, GEE focuses on the average trend. In the lab it is preferred for
    population-level questions like "what is the time effect in the average sample?".

====================================================================================

#56  GLMM
    file: 56_glmm_type_count.xlsx
  >> SCENARIO (narration):
    We model a repeatedly measured type count at the same sites with time and
    intervention, taking site as a random effect. For repeated counts, GLMM is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: type_count
    - Predictor(s): time
    - Predictor(s): intervention
    - Cluster: site_id (Poisson)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    intervention coef = -0.661   p < .001 ***   outcome: type_count (count, Poisson)   group: site_id

>> COMMENTARY (narration):
    We modeled a COUNT outcome (type count) measured repeatedly at the same sites with time and intervention, taking
    site as a random effect; the intervention effect is significant and negative (b = -0.66, p < .001). GLMM extends
    LMM to non-normal outcomes (count, binary): it handles both the distribution (Poisson) and the clustering (random
    effect) at the same time. In the lab/ecology it is the right model for repeatedly measured count outcomes
    (site-wise species count).

====================================================================================

#57  Elastic Net
    file: 57_elasticnet.xlsx
  >> SCENARIO (narration):
    We model a feature from 40 predictors with Elastic Net. For many clustered
    predictors, Elastic Net is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: feature
    - Predictor(s): x01..40

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.097   L1 ratio = 0.50   R^2 (train) = 0.538   40 predictors

>> COMMENTARY (narration):
    We modeled a feature from 40 predictors with Elastic Net. Elastic Net blends the ridge (L2) and lasso (L1)
    penalties (L1 ratio = 0.5): it both keeps groups of correlated variables together (ridge property) and zeroes out
    redundant ones (lasso property). It thus provides a balanced model with high-dimensional, clustered predictors. In
    the lab it is chosen when there are many clustered measurements, where lasso or ridge alone is insufficient.

====================================================================================

#58  Robust Regression
    file: 58_robust_outlier.xlsx
  >> SCENARIO (narration):
    We model density with temperature, down-weighting outliers. For data with
    outliers, robust regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: density_g_cm3
    - Predictor(s): temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust Intercept = 1.010   OLS Intercept = 0.997   outcome: density_g_cm3   predictor: temperature_C

>> COMMENTARY (narration):
    When modeling density with temperature, we used robust regression to prevent outliers from distorting the estimate.
    The robust and OLS intercepts are close (1.010 vs 0.997), indicating outliers have limited influence in this data;
    still, the robust method secures the "typical" relationship by down-weighting extreme observations. In the lab it
    is the right way to get robust estimates without deleting outliers (measurement error/extreme values) in data that
    contains them.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_yield.xlsx
  >> SCENARIO (narration):
    We model different points of yield's distribution (lower 10%, median, upper 90%)
    separately. For varying effects, quantile regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: yield
    - Predictor(s): concentration
    - Predictor(s): temperature_C
    - Quantiles: 0.10 / 0.50 / 0.90

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    separate slopes reported for q=0.10, q=0.50, q=0.90   outcome: yield   predictors: concentration, temperature_C

>> COMMENTARY (narration):
    We modeled not just yield's mean but different points of the distribution (lower 10%, median, upper 90%) separately.
    Predictor effects can vary by quantile -- a factor may be strong under low-yield conditions and weak under
    high-yield ones. Quantile regression gives the true picture when the "mean effect" is misleading (effect varies
    across the distribution). In the lab it shows what classic regression misses by examining extreme-condition
    behavior and yield variability.

====================================================================================

#60  ROC Curve
    file: 60_roc_biomarker.xlsx
  >> SCENARIO (narration):
    We assess how well a biomarker separates a binary outcome with ROC. For
    discrimination and threshold selection, ROC is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: positive_label
    - Predictor(s): biomarker

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.964 (excellent discrimination)   Youden optimum threshold = 0.621   Sensitivity = 0.920, 1-Specificity = 0.136
    n = 250 (positive 162, negative 88)

>> COMMENTARY (narration):
    We assessed how well a biomarker separates a binary outcome (positive_label) with a ROC curve: AUC = 0.96,
    excellent discrimination; the optimum decision threshold was set at 0.62 via Youden. ROC shows the
    sensitivity-specificity trade-off at all possible thresholds and evaluates the model without being tied to a single
    threshold. In the lab it is the standard tool for measuring a diagnostic/biomarker test's discriminative power and
    selecting the best decision threshold.

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss_type_distribution.xlsx
  >> SCENARIO (narration):
    We measure a classification/distribution model's discrimination with TSS. For a
    fair measure under class imbalance, TSS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: asset
    - Predictor(s): feature_1
    - Predictor(s): feature_2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.35 (acceptable)   Sensitivity = 0.53   Specificity = 0.82   N = 200
    binary outcome: asset   predictors: feature_1, feature_2

>> COMMENTARY (narration):
    We measured a classification/distribution model's discrimination with TSS: TSS = 0.35 (sensitivity 0.53,
    specificity 0.82). TSS = sensitivity + specificity - 1; it excludes chance-expected success and gives a fair
    performance measure even with imbalanced classes. TSS is a common metric in species distribution modeling (SDM) and
    ecology. While accuracy can be biased, TSS evaluates the power to capture both presence and absence together. In
    the lab/ecology it is robust for reporting the true discriminative power of classification models.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity_3sinif.xlsx
  >> SCENARIO (narration):
    We evaluate a classifier's predictions against true labels. For detailed
    performance, the confusion matrix is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: actual_label
    - 2nd measure / group: prediction_label

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.93   F1 = 0.918   (true label vs predicted label)

>> COMMENTARY (narration):
    We evaluated a classifier's predictions against true labels via a confusion matrix: accuracy 93%, F1 = 0.92. The
    confusion matrix gathers true/false positives and negatives in one table, from which sensitivity, specificity,
    precision and F1 are derived. Because a single accuracy number can mislead (especially with imbalanced classes),
    this metric set shows where the model errs. In the lab it is fundamental for detailed reporting of
    prediction/classification model performance.

====================================================================================

#63  Random Forest
    file: 63_rf_class.xlsx
  >> SCENARIO (narration):
    We classify class from four features with a random forest. For nonlinear
    prediction, random forest is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: class
    - Predictor(s): feature_a
    - Predictor(s): feature_b
    - Predictor(s): feature_c
    - Predictor(s): feature_d

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.986   outcome: class   predictors: feature_a, feature_b, feature_c, feature_d

>> COMMENTARY (narration):
    We classified class from four features with a random forest: accuracy 98.6%. Random forest combines the votes of
    hundreds of decision trees; it automatically captures nonlinear relationships and interactions and also gives a
    variable-importance ranking. It reduces a single tree's overfitting by averaging. In the lab it is widely used for
    complex, nonlinear classification problems (type/phase prediction) because it offers both high accuracy and "which
    variable matters" information.

====================================================================================

#64  Support Vector Machines (SVM)
    file: 64_svm_binary.xlsx
  >> SCENARIO (narration):
    We classify the binary label from four features with SVM. For well-separable
    classes, SVM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: label
    - Predictor(s): x1
    - Predictor(s): x2
    - Predictor(s): x3
    - Predictor(s): x4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: label   predictors: x1, x2, x3, x4

>> COMMENTARY (narration):
    We classified the binary label from four features with SVM: accuracy 100% -- the classes are perfectly separable
    with these features (very high accuracy should be confirmed against overfitting via cross-validation). SVM finds
    the decision boundary separating classes with the widest margin; with the kernel trick it can also do nonlinear
    separation. In the lab it is a strong, stable classifier for well-separable class problems, especially on
    small-to-medium data.

====================================================================================

#65  Gradient Boosting
    file: 65_gradient_boosting.xlsx
  >> SCENARIO (narration):
    We classify the risk class from four predictors with gradient boosting. For top
    accuracy, gradient boosting is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: risk
    - Predictor(s): x1
    - Predictor(s): x2
    - Predictor(s): x3
    - Predictor(s): x4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.875   outcome: risk   predictors: x1, x2, x3, x4

>> COMMENTARY (narration):
    We classified the risk class from four predictors with gradient boosting: accuracy 87.5%. Gradient boosting adds
    weak trees sequentially -- each new tree corrects the previous model's errors; this is why it wins most prediction
    competitions. While random forest votes in parallel, boosting reduces error step by step. In the lab it is
    preferred where the highest predictive accuracy is sought (risk classification, outcome prediction); it requires
    tuning against overfitting.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_4kume.xlsx
  >> SCENARIO (narration):
    We cluster samples by two features. For sample grouping, k-means is appropriate.
  >> VARIABLE SELECTION:
    - Variables: feature_a
    - Variables: feature_b
    - Number of clusters: 4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.627   variables: feature_a, feature_b

>> COMMENTARY (narration):
    We split samples into 4 clusters by two features: silhouette = 0.63, a good separation. K-means assigns
    observations to the nearest cluster center and iteratively updates the centers; it groups similar samples into
    natural clusters. No labels are needed (unsupervised). In the lab it is the core method for grouping
    samples/specimens and discovering patterns; silhouette checks the appropriateness of the cluster count.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5tur.xlsx
  >> SCENARIO (narration):
    We hierarchically cluster samples by ten features. For nested group structure,
    hierarchical clustering is appropriate.
  >> VARIABLE SELECTION:
    - Variables: feat_01..10
    - Number of clusters: 5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 5   Silhouette = 0.136   variables: feat_01..10 (10 features)

>> COMMENTARY (narration):
    We split samples into 5 groups by ten features with hierarchical clustering (silhouette = 0.14, weak separation --
    groups partly overlap). Hierarchical clustering merges observations step by step to build a tree (dendrogram); its
    difference from k-means is not having to fix the cluster count in advance and seeing the nested structure. In the
    lab it is used to explore a hierarchy of types/groups (main group -> sub-group) and to read the natural cluster
    count from the dendrogram.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan_location.xlsx
  >> SCENARIO (narration):
    We cluster sample locations with density-based DBSCAN. For spatial clusters and
    outliers, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Variables: lat
    - Variables: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 5   Silhouette = 0.887   variables: lat, lon (location)

>> COMMENTARY (narration):
    We clustered sample locations (lat, lon) with density-based DBSCAN: 5 dense clusters, silhouette = 0.89, very good
    separation. Unlike k-means, DBSCAN does not require the cluster count in advance, can find clusters of any shape,
    and marks sparse points as "noise". In the lab/field work it is ideal for detecting spatial concentrations (sample
    clusters, temperature zones) and isolating outlier locations.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca_6ozellik.xlsx
  >> SCENARIO (narration):
    We reduce six correlated features to a few components with PCA. For
    dimensionality reduction, PCA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: f1
    - Variables: f2
    - Variables: f3
    - Variables: f4
    - Variables: f5
    - Variables: f6

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 explains 92.7% of variance   variables: f1..f6 (6 features)

>> COMMENTARY (narration):
    We reduced six correlated features to a few components with PCA: the first component alone explains 92.7% of
    variance -- so these six measures largely reflect a single latent dimension. PCA transforms correlated variables
    into mutually independent components; it reduces dimensions, eases visualization and resolves multicollinearity. In
    the lab it is fundamental for summarizing multi-indicator measurements and building a "composite index".

====================================================================================

#70  t-SNE
    file: 70_tsne_5tip.xlsx
  >> SCENARIO (narration):
    We reduce 12-sensor-dimensional data to 2 dimensions with t-SNE for
    visualization. For nonlinear cluster discovery, t-SNE is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sensor_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence = 0.506 (good)   high-dimensional 12 sensors embedded into 2 dimensions

>> COMMENTARY (narration):
    We reduced 12-sensor-dimensional data to 2 dimensions for visualization with t-SNE (KL = 0.51, good quality). t-SNE
    tries to preserve high-dimensional neighborhoods, placing similar observations near and dissimilar ones far; it
    reveals nonlinear cluster structures visually that PCA misses. Interpretation is visual (the axes have no absolute
    meaning). In the lab it is used for exploratory visualization of hidden type clusters in high-dimensional
    sensor/measurement data.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds_3grup.xlsx
  >> SCENARIO (narration):
    We place the inter-sample similarity structure onto a 2D map. For similarity
    maps, MDS is appropriate.
  >> VARIABLE SELECTION:
    - Variables: f1
    - Variables: f2
    - Variables: f3
    - Variables: f4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.053 (acceptable)   variables: f1..f4

>> COMMENTARY (narration):
    We placed the distance/similarity structure among samples onto a 2-dimensional map with low stress (0.05): low
    stress means the map represents the true distances well. MDS positions observations in an interpretable space by
    preserving inter-observation distances; its difference from t-SNE is the aim of preserving global distance
    structure. In the lab it is used to draw sample/group similarity maps and for positioning between groups.

====================================================================================

#72  UMAP
    file: 72_umap_5tip.xlsx
  >> SCENARIO (narration):
    We reduce 25-dimensional feature data to 2 dimensions with UMAP. For both local
    and global structure, UMAP is appropriate.
  >> VARIABLE SELECTION:
    - Variables: feat_01..25

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    25 features embedded into 2 dimensions   (local-global balance via n_neighbors / min_dist)

>> COMMENTARY (narration):
    We reduced 25-dimensional feature data to 2 dimensions with UMAP. UMAP does nonlinear dimensionality reduction like
    t-SNE but preserves both local and global structure better and is faster. Larger n_neighbors emphasizes broader
    groups, larger min_dist emphasizes the gaps between clusters. In the lab it is a modern choice for visualizing
    high-dimensional measurement data and exploring natural cluster structure.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_21madde.xlsx
  >> SCENARIO (narration):
    We measure the internal consistency of a 21-item scale. For scale reliability,
    Cronbach's alpha is appropriate.
  >> VARIABLE SELECTION:
    - Variables: item_01..21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach alpha = 0.967 (excellent internal consistency)   21 items

>> COMMENTARY (narration):
    We measured the internal consistency of a 21-item scale with Cronbach's alpha: alpha = 0.97, excellent -- the items
    measure the same construct consistently. Cronbach's alpha shows how much the items of a scale "move together";
    above 0.70 is considered acceptable. (A very high value can also signal item redundancy.) In the lab it is the
    standard index for reporting the reliability of rating/evaluation scales.

====================================================================================

#74  Likert Scale Analysis
    file: 74_likert_3boyut.xlsx
  >> SCENARIO (narration):
    We analyze a 15-item Likert set of three dimensions. For ordinal scale summary
    and reliability, Likert analysis is appropriate.
  >> VARIABLE SELECTION:
    - Variables: dimension1_1..5 / dimension2_1..5 / dimension3_1..5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of items k = 15   Cronbach alpha = 0.820 (good)

>> COMMENTARY (narration):
    We analyzed a 15-item Likert set of three dimensions: reliability is good (alpha = 0.82), and item distributions
    and central tendencies were reported. Likert analysis describes ordinal scale responses (1-5) with appropriate
    summaries (median, distribution, pile-up) and checks scale reliability. In the lab it is used for correctly
    summarizing and interpreting evaluation/perception scales.

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18madde.xlsx
  >> SCENARIO (narration):
    We discover the latent factors behind eighteen items. For scale-structure
    discovery, EFA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: q01..18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO = 0.907 (excellent)   factor analysis appropriate   18 items (q01..18)

>> COMMENTARY (narration):
    We applied EFA to discover how many latent factors lie behind eighteen items: KMO = 0.91, the data is very suitable
    for factor analysis. EFA reduces observed items to a few unobserved "factors"; it reveals which items measure the
    same dimension. In the lab it is fundamental when developing a new scale (discovering structure) and mapping item
    groups to theoretical dimensions; KMO and Bartlett are prerequisite tests.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3olcum.xlsx
  >> SCENARIO (narration):
    We measure the consistency of three repeated measurements. For repeat
    reliability on continuous measures, ICC is appropriate.
  >> VARIABLE SELECTION:
    - Variables: measurement_a
    - Variables: measurement_b
    - Variables: measurement_c

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.938 (Excellent)   95% CI [0.900, 0.960]   3 measurements (measurement_a/b/c)

>> COMMENTARY (narration):
    We measured the consistency of three repeated measurements (measurement_a/b/c) with ICC: ICC = 0.94, excellent --
    the measurements are nearly identical. Unlike kappa, ICC measures reliability across observers/repeats on
    CONTINUOUS measures and can assess both consistency and absolute agreement. In the lab it is the standard index for
    determining how consistent/interchangeable a sample's repeated measurements (instrument repeatability, observer
    reliability) are.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_12madde_3faktor.xlsx
  >> SCENARIO (narration):
    We test whether a predefined three-factor structure fits the data. For construct
    validity, CFA is appropriate.
  >> VARIABLE SELECTION:
    - Value: dimension1: dimension1_1..4
    - Value: dimension2: dimension2_1..4
    - Value: dimension3: dimension3_1..4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.993   RMSEA = 0.027   3 factors: dimension1/2/3 (4 items each)

>> COMMENTARY (narration):
    Unlike EFA, we TESTED whether a pre-defined three-factor structure (3 dimensions) fits the data with CFA: the fit
    indices are very good (CFI = 0.99, RMSEA = 0.03). CFA tests a theoretical scale model -- which item loads on which
    factor is fixed in advance, and the question is "does the model fit the data?". In the lab it is a mandatory step
    for confirming the construct validity of a developed scale/measurement model.

====================================================================================

#78  Survey Mean
    file: 78_survey_means.xlsx
  >> SCENARIO (narration):
    We estimate the measurement mean in a stratified/weighted sample. For a complex
    sample, design-based mean is appropriate.
  >> VARIABLE SELECTION:
    - Variables: measurement
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    measurement: M_hat = 45.35   SE = 0.42   95% CI [44.52, 46.17]   CV 0.92%
    weight: weight, stratum: region

>> COMMENTARY (narration):
    In a stratified/weighted sample design we estimated the measurement mean while accounting for the design: M = 45.35,
    with a Taylor-linearization SE and 95% CI [44.5, 46.2]. Complex-sample methods account for unequal selection
    probabilities (weights) and stratification; ignoring these biases the standard errors. In the lab/field studies it
    is the right way to produce correct point estimates and confidence intervals from stratified, weighted sampling.

====================================================================================

#79  Survey Frequency
    file: 79_survey_freq.xlsx
  >> SCENARIO (narration):
    We estimate category proportions of a categorical variable accounting for the
    design. For a complex sample, design-based frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: category
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    weighted proportions + Taylor SE + 95% CI for category categories
    (e.g. 'low' p_hat = 0.281, 'high' p_hat = 0.077)

>> COMMENTARY (narration):
    We estimated the category proportions of a categorical variable (category) under sampling weights and
    stratification; each proportion has a design-based SE and confidence interval. Complex-sample frequency analysis,
    unlike a simple percentage, estimates population proportions without bias by accounting for the sampling design. In
    the lab/field studies it is used to correctly report category distributions from stratified samples.

====================================================================================

#80  Survey Total
    file: 80_survey_total.xlsx
  >> SCENARIO (narration):
    We estimate the population total from the sample. For a complex sample,
    design-based total is appropriate.
  >> VARIABLE SELECTION:
    - Variables: amount
    - Weight: weight
    - Stratum: stratum

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    amount: T_hat = 266,166   SE = 3,389.61   95% CI [259,493, 272,839]   stratum: stratum

>> COMMENTARY (narration):
    We estimated the population TOTAL (total amount) from the sample: T = 266,166, 95% CI [259,493, 272,839]. Weights
    tell how many population units each observation represents; the total is estimated by summing those weights, with
    uncertainty reported via Taylor SE. In the lab/field studies it is the right method for producing population-scaled
    totals (total stock, total amount) from a sample.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg.xlsx
  >> SCENARIO (narration):
    We regress an outcome accounting for the design. For relationships in complex
    samples, design-based regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: result
    - Predictor(s): age
    - Predictor(s): education
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.752   age coef = 79.86 (p < .001)   outcome: result   predictors: age, education   weight: weight

>> COMMENTARY (narration):
    We regressed an outcome while accounting for the sampling design (weight, stratum): age is a significant predictor
    (b = 79.86, p < .001), and the model explains 75% of variance. Design-based regression incorporates weights and the
    cluster/stratum structure into coefficient and standard-error calculation; ordinary regression ignores these and
    gives biased inference. In the lab/field studies it is the right way to model relationships in representative sample
    data.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic.xlsx
  >> SCENARIO (narration):
    We model a binary survey outcome with design-weighted logistic. For binary
    outcomes in complex samples, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: binary_result
    - Predictor(s): age
    - Predictor(s): measurement
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    age: OR = 0.995   p = 0.540 ns   binary outcome: binary_result   predictors: age, measurement   weight: weight

>> COMMENTARY (narration):
    We modeled a binary outcome (binary_result) with design-weighted logistic regression: the age effect is
    non-significant (OR = 0.995, p = 0.54). This method extends logistic regression to complex sample designs -- weight
    and stratum are reflected in the standard errors. Here the "no significant effect" result is also valuable. In the
    lab/field studies it is used to correctly estimate the probability of a binary outcome from representative samples.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_measurement_temperature.xlsx
  >> SCENARIO (narration):
    We model a measurement with temperature via a flexible curve. For a nonlinear
    effect, GAM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: measurement
    - Predictor(s): temperature_C (smooth)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 (explained) = 0.609   outcome: measurement   predictor: temperature_C (smooth term)

>> COMMENTARY (narration):
    We modeled a measurement with temperature, without assuming a straight line in advance, via a flexible curve
    (smooth): the explained variance is high (pseudo R^2 = 0.61). GAM extends linear regression -- it models each
    predictor's effect as a smooth function learned from the data, capturing curved relationships without losing
    interpretability. In the lab it is a more explanatory choice than black-box models for relationships where the
    effect is nonlinear (saturation, threshold, optimum).

====================================================================================

#84  Discriminant Analysis
    file: 84_diskriminant_3sinif.xlsx
  >> SCENARIO (narration):
    We classify class from five continuous measures. To assign to predefined
    classes, discriminant analysis is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: class
    - Predictor(s): f1
    - Predictor(s): f2
    - Predictor(s): f3
    - Predictor(s): f4
    - Predictor(s): f5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: class   predictors: f1, f2, f3, f4, f5

>> COMMENTARY (narration):
    We classified class from five continuous measures with discriminant analysis: accuracy 100% -- the classes are
    fully separable with these measures (overfitting risk should be checked via cross-validation). Discriminant
    analysis finds the linear combinations that best separate groups; it both classifies and shows which variable is
    most influential in separation. In the lab it is used to assign new samples to predefined classes and to identify
    discriminating features.

====================================================================================

#85  Conditional Logit
    file: 85_conditional_logit.xlsx
  >> SCENARIO (narration):
    We model participants' choices among alternatives by option features. For
    discrete-choice data, conditional logit is appropriate.
  >> VARIABLE SELECTION:
    - Chooser: participant_id
    - Choice: chosen
    - Alternative features: price
    - Alternative features: quality

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McFadden pseudo R^2 = 0.516   chooser: participant_id   choice: chosen   features: price, quality

>> COMMENTARY (narration):
    We modeled the choices participants made among alternatives by the alternatives' features (price, quality) with a
    conditional logit (pseudo R^2 = 0.52). This model is for "discrete choice" data where each individual picks one
    from a choice set; it estimates how an option's features affect its probability of being chosen. Its difference
    from standard logistic is that the choice is conditional on the individual's option set. It is the core method for
    preference/choice experiments.

====================================================================================

#86  Kaplan-Meier Survival
    file: 86_km_survival.xlsx
  >> SCENARIO (narration):
    We examine the time to an event and group differences. For censored time data,
    Kaplan-Meier is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Event: event
    - Grouping (categorical): category

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 8.36   Log-rank chi^2(1) = 0.013, p = 0.911 ns   time: duration_year, event: event, group: category

>> COMMENTARY (narration):
    We examined the time to an event with Kaplan-Meier: median survival 8.36 units; group (category) survival curves
    were compared with the log-rank test (p = 0.91, no difference). KM correctly handles censored time data (those
    whose event has not yet occurred); a plain average ignores these observations and is biased. In the lab/biology it
    is fundamental for "lifetime", durability and degradation-timing analysis.

====================================================================================

#87  Cox Proportional Hazards
    file: 87_cox_event.xlsx
  >> SCENARIO (narration):
    We examine event risk with continuous predictors in a Cox model. For multiple
    predictors in censored time, Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Event: event
    - Predictor(s): age
    - Predictor(s): risk_1
    - Predictor(s): risk_2
    - Predictor(s): risk_3

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.577   predictors (HR): age=1.004 (p=0.53), risk_1=0.995 (p=0.86), risk_2=0.777 (p=0.14), risk_3=1.007 (p=0.12)
    time: duration_year, event: event   (none significant alone)

>> COMMENTARY (narration):
    We examined event risk with continuous predictors (age, risk_1..3) in a Cox model: in this sample no predictor was
    significant alone (closest risk_2: HR = 0.78, p = 0.14), concordance 0.58 (weak-moderate discrimination). Cox
    regression gives the effect of several predictors on "event time" as hazard ratios (HR) in censored time data,
    making no assumption about the baseline hazard's shape. In the lab/biology it is the gold standard for identifying
    the factors that drive event risk; here no decisive factor was found.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT)
    file: 88_aft_weibull.xlsx
  >> SCENARIO (narration):
    We model survival time assuming a Weibull distribution. For explicit time
    estimation, AFT is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Event: event
    - Predictor(s): age_covariate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution: Weibull   Median survival = 12.25   time: duration_year, event: event, predictor: age_covariate

>> COMMENTARY (narration):
    We modeled survival time by assuming a parametric (Weibull) distribution: median survival ~12.3 units. AFT
    (accelerated failure time) models, unlike Cox, choose an explicit distribution for the hazard shape and directly
    interpret how predictors "accelerate/decelerate" time. If the data fit the assumed distribution they are more
    powerful than Cox. In the lab it is preferred when explicit time estimation and extrapolation are needed.

====================================================================================

#89  Competing Risks
    file: 89_competing_risks.xlsx
  >> SCENARIO (narration):
    We model a setting where several distinct ends are possible. For mutually
    exclusive events, competing risks is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Event: event_cause -> olay tipi / event type
    - Predictor(s): age_covariate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 60   different event causes (event_cause) modeled as distinct event types

>> COMMENTARY (narration):
    We examined a setting where several distinct ends are possible in the competing-risks framework: 60 observations
    censored, the rest split into distinct event types. While standard survival treats all events alike, the
    competing-risks method accounts for the fact that "once one occurs the others no longer can" and gives a separate
    cumulative incidence for each event type. In the lab/biology it is necessary to correctly model a unit's mutually
    exclusive distinct ends (different degradation/death causes).

====================================================================================

#90  Time-Dependent Cox
    file: 90_tvcox.xlsx
  >> SCENARIO (narration):
    We build a Cox model where the predictor changes over time. For a time-varying
    covariate, time-dependent Cox is appropriate.
  >> VARIABLE SELECTION:
    - Unit (id): sample_id
    - Start: start
    - Stop: end
    - Event: event
    - Predictor(s): measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    measurement: HR = 1.015   p = 0.300 ns   (time-varying covariate)   id: sample_id, start/stop intervals

>> COMMENTARY (narration):
    We built a Cox model where the predictor (measurement) CHANGES over time: for each unit the current measurement
    value was used via start-stop intervals; the effect here is non-significant (HR = 1.015, p = 0.30). Time-dependent
    Cox correctly handles non-constant covariates (changing measurement, changing condition) -- it uses the predictor's
    current value at the event time. In the lab it is the right method for modeling the effect of time-varying risk
    factors on an event.

====================================================================================

#91  Survey Cox Regression
    file: 91_survey_phreg.xlsx
  >> SCENARIO (narration):
    We carry survival analysis into a complex survey design (weight+cluster). For
    design-faithful event time, survey_phreg is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: event
    - Predictor(s): region (faktorize/factorized)
    - Weight: weight
    - Cluster: cluster_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.51   region_n: p = 0.696 ns   time: duration_month, event: event   weight: weight, cluster: cluster_id

>> COMMENTARY (narration):
    We carried survival analysis into a complex survey design (weight + cluster) with a Cox model: the region effect is
    non-significant (p = 0.70), concordance 0.51 (weak discrimination). survey_phreg extends the proportional-hazards
    model to a stratum/cluster/weight structure -- standard errors are corrected for the design. In the lab/field
    studies it is the design-faithful way to model event time in representative panel/sample data.

====================================================================================

#92  Interval-Censored Survival
    file: 92_interval_censored.xlsx
  >> SCENARIO (narration):
    We model data where the event is known only within an interval. For interval
    censoring, this is appropriate.
  >> VARIABLE SELECTION:
    - Lower bound: left_censor
    - Upper bound: survival_censor

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of events = 29   Median survival = 10.0   (lower bound: left_censor, upper bound: survival_censor)

>> COMMENTARY (narration):
    We modeled data where the exact event time is unknown and only known to have occurred within an INTERVAL: median
    survival 10 units, 29 events. Interval-censored methods are for cases where the event lies "somewhere between two
    observations" (between periodic measurements); fixing the event to the interval's mid/end point biases results,
    while this method carries the uncertainty correctly. In the lab it is used to correctly model events occurring
    between periodic inspections (deterioration between two measurements).

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty_cox.xlsx
  >> SCENARIO (narration):
    We model the group-specific hidden risk as frailty in clustered survival data.
    For shared hidden risk, frailty Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Event: event
    - Predictor(s): clinical_group (faktorize/factorized)
    - Cluster: group_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.556   cg_n: p = 0.236 ns   time: duration_year, event: event   cluster: group_id (frailty)

>> COMMENTARY (narration):
    In clustered survival data (units within the same group) we modeled the group-specific unobserved risk as
    "frailty" (a random effect); the clinical-group effect is non-significant (p = 0.24). Frailty Cox accounts for the
    hidden risk shared by units in the same cluster -- solving the independence assumption that standard Cox violates.
    In the lab/biology it is the right choice for event data clustered within a batch/group (shared hidden risk).

====================================================================================

#94  Time Series Analysis
    file: 94_ts_monthly.xlsx
  >> SCENARIO (narration):
    We examine a monthly measurement series (trend, season, stationarity). For
    time-dependent structure, time series analysis is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.866 (not stationary)   Trend: increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly measurement series: the ADF test shows it is not stationary (p = 0.87), with an increasing
    trend and seasonality. Time series analysis reveals the time-dependent structure (trend, season, autocorrelation)
    of observations; ordinary statistics are misleading because of this dependency. Stationarity is a prerequisite for
    models like ARIMA; if non-stationary, differencing is needed. In the lab/environmental monitoring it is the
    starting step for analyzing time-series measurements.

====================================================================================

#95  STL Decomposition
    file: 95_stl_daily.xlsx
  >> SCENARIO (narration):
    We decompose the measurement series into trend, season and residual. For
    seasonal decomposition, STL is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Season period = 7   series decomposed into trend + season + residual components

>> COMMENTARY (narration):
    We decomposed the measurement series into three components with STL: trend, the seasonal pattern (period = 7,
    weekly), and residual. STL visually and numerically separates a time series' long-term trend, recurring seasonal
    pattern and unexplained fluctuation; this answers "what is the underlying trend, and how much is seasonal?". In the
    lab/environmental monitoring it is fundamental for de-seasonalizing series to see the underlying trend and for
    anomaly detection.

====================================================================================

#96  ARIMA Forecast
    file: 96_arima_monthly.xlsx
  >> SCENARIO (narration):
    We model the measurement series with ARIMA and produce a forecast. For series
    forecasting, ARIMA is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model selected   AIC = 2450.18   forecast + confidence interval

>> COMMENTARY (narration):
    We modeled the series with ARIMA and produced a forecast (AIC = 2450.18 for model selection). ARIMA forecasts the
    future from the series' own past values (AR), trend (I - differencing) and past errors (MA); on a stationarized
    series it gives strong short-to-medium-term forecasts. In the lab/environmental monitoring it is the most common
    classical method for measurement-series forecasting; the confidence interval shows the forecast uncertainty.

====================================================================================

#97  Exponential Smoothing (ETS)
    file: 97_ets_monthly.xlsx
  >> SCENARIO (narration):
    We model the measurement series with Holt-Winters exponential smoothing. For a
    seasonal trended series, ETS is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method: Holt-Winters (seasonal additive, period = 12)   AIC = 830.58

>> COMMENTARY (narration):
    We modeled the measurement series with Holt-Winters exponential smoothing: it jointly estimates the level, trend
    and a 12-period seasonal component. ETS tracks the series' current level, trend and season by weighting recent
    observations more (exponentially decaying weights); for seasonal and trended series it is a practical alternative
    to ARIMA. In the lab/environmental monitoring it gives fast, reliable results for forecasting series with regular
    seasonal patterns.

====================================================================================

#98  Mann-Kendall Trend
    file: 98_mann_kendall_trend.xlsx
  >> SCENARIO (narration):
    We test a significant trend in the measurement series without distributional
    assumptions. For robust trend detection, Mann-Kendall is appropriate.
  >> VARIABLE SELECTION:
    - Value: measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   significant trend detected   (trend magnitude via Sen's slope)   variable: measurement

>> COMMENTARY (narration):
    We tested whether the measurement series has a significant trend -- without assuming a distribution -- with
    Mann-Kendall: p < .001, a significant trend; Sen's slope gives the robust (outlier-resistant) magnitude of the
    trend. Mann-Kendall is a nonparametric trend test; it requires no normality and is resistant to outliers, hence
    common in environmental and time-series work. In the lab/environmental monitoring it is used to robustly detect
    long-term measurement trends (increasing/decreasing).

====================================================================================

#99  Anomaly Detection
    file: 99_anomali_measurements.xlsx
  >> SCENARIO (narration):
    We detect unusual observations in multivariate data. For composite outlier
    detection, anomaly detection is appropriate.
  >> VARIABLE SELECTION:
    - Variables: f1
    - Variables: f2
    - Variables: f3
    - Variables: f4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 11   variables: f1, f2, f3, f4

>> COMMENTARY (narration):
    We automatically detected unusual observations in multivariate data: 11 samples were flagged as anomalies. Anomaly
    detection finds observations that deviate markedly from normal (erroneous record, contamination, exceptional case)
    by evaluating several variables together; it catches "composite" outliers that univariate thresholds miss. In the
    lab it is used for quality control, erroneous-measurement detection and the early detection of exceptional sample
    behavior.

====================================================================================

#100  Variance Components
    file: 100_varcomp_3seviye_h2.xlsx
  >> SCENARIO (narration):
    We decompose measurement variability into nested levels (upper/lower unit). For
    hierarchical variability, variance components is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): upper_unit
    - Factor (categorical): lower_unit (nested)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Random factors: upper_unit + lower_unit (nested within upper)   % contribution of each component + residual

>> COMMENTARY (narration):
    We decomposed a measurement's variability into nested levels -- upper unit and lower unit within upper: the percent
    contribution of each level and the residual were reported. Variance components analysis answers "how much of the
    variability is between upper units, how much between lower units, how much within unit?". In the lab it is used in
    hierarchical structures (batch>sample>replicate) to see where uncertainty concentrates and in sampling/measurement
    design; it underlies ratio estimates such as heritability (h2).

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old.xlsx
  >> SCENARIO (narration):
    We examine two groups' measurement difference with a Bayesian t-test, expressing
    evidence as a Bayes factor. For an intuitive evidence ratio, this is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Grouping (categorical): group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 3.36e+46   Cohen's d = 3.939   (overwhelming evidence for H1)   outcome: value, group: group

>> COMMENTARY (narration):
    We examined the measurement difference between two groups (new/old method) with a Bayesian t-test: the Bayes factor
    is overwhelming (BF10 ~ 3.4e46), the effect very large (d = 3.94) -- the data support the "difference" hypothesis
    over "no difference" by astronomical odds. Unlike a p-value, the Bayes factor gives the RELATIVE evidence strength
    of two hypotheses and can distinguish "no evidence" from "no difference". In the lab it is preferred when one wants
    to express the evidential strength of a decision as an intuitive ratio.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BF10.xlsx
  >> SCENARIO (narration):
    We evaluate the relationship between two variables in a Bayesian framework. For
    evidential strength, Bayesian correlation is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: X_variable
    - 2nd measure / group: Y_variable

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.780 (very strong)   BF10 = 4.50e+14   (overwhelming evidence for a relationship)

>> COMMENTARY (narration):
    We evaluated the relationship between two variables in a Bayesian framework: r = 0.78 and BF10 ~ 4.5e14, i.e. very
    strong evidence for a relationship. Bayesian correlation, instead of a classical p-value, presents the evidential
    strength of the relationship as a Bayes factor and the coefficient's posterior distribution. In the lab it is
    valuable for reporting not just whether the relationship between two measurements is "significant" but how strongly
    it is "evidenced".

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_2yonlu.xlsx
  >> SCENARIO (narration):
    We examine a measurement's difference across two factors with Bayesian ANOVA.
    For the evidential strength of factor effects, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): factor1
    - 2nd Factor: factor2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor1: BF10 = 76099 (very strong evidence)   outcome: value, factors: factor1, factor2

>> COMMENTARY (narration):
    We examined a measurement's difference across two factors with Bayesian ANOVA: for factor1, BF10 ~ 76,000, very
    strong evidence. Bayesian ANOVA compares the effects of factors and their interactions via Bayes factors; it ranks
    probabilistically which model (which effects) best explains the data. Unlike classical ANOVA's "reject/don't
    reject" decision, it quantifies the relative support among models. In the lab it is used to compare the evidential
    strength of factor effects.

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_LMM.xlsx
  >> SCENARIO (narration):
    We analyze group-nested data with a Bayesian hierarchical model. For stable
    estimates in small groups, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_response
    - Cluster: group_id
    - Predictor(s): X_covariate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2_u (group) = 2.90 (8.6%)   sigma^2_eps (residual) = 30.79 (91.4%)   outcome: Y_response, group: group_id

>> COMMENTARY (narration):
    We analyzed group-nested data with a Bayesian hierarchical (multilevel) model: 8.6% of variability is
    between-group, 91.4% within-group (ICC ~ 0.09). The Bayesian hierarchical model is the Bayesian version of LMM -- it
    estimates group effects with posterior distributions and balances small groups via "partial pooling". In the lab it
    is powerful for producing stable estimates even in small groups within multilevel (batch/sample/replicate) data.

====================================================================================

#105  Spatial SAR
    file: 105_spatial_sar_spatial.xlsx
  >> SCENARIO (narration):
    When modeling a measurement we handle spatial spillover with SAR. For
    neighborhood effects, SAR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho = 0.336   z = 4.94   p < .001 ***   Pseudo R^2 = 0.87   (N = 100, k-NN W, k = 5)
    outcome: Y_value, predictors: X1, X2

>> COMMENTARY (narration):
    When modeling a measurement (Y_value), we handled spatial spillover (the effect of neighboring units) with SAR: the
    spatial lag parameter is significant and positive (rho = 0.34, p < .001) -- a unit's value is related to its
    neighbors' value, a "cluster/spillover" pattern. SAR incorporates spatial dependency into the model; if ignored,
    standard errors are biased. In the lab/environmental studies it is the right method for modeling the geographic
    spread of measurements (neighborhood effect).

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_residual.xlsx
  >> SCENARIO (narration):
    We model spatial dependency in the error term. For unmeasured geographic
    factors, SEM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda = 0.590   z = 6.38   p < .001 ***   Pseudo R^2 = 0.84   (N = 100, k-NN W, k = 5)

>> COMMENTARY (narration):
    This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
    0.59, p < .001) -- the effect of geographic variables omitted from the model makes neighboring errors correlated.
    Unlike SAR, SEM attributes the spread to the error rather than the outcome. In the lab/environmental studies, when
    the source of spatial autocorrelation is unmeasured geographic factors (climate, soil), the correct specification
    is SEM; it is chosen by comparison with SAR.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local.xlsx
  >> SCENARIO (narration):
    Assuming the relationship is not constant in space, we estimate separate
    coefficients per location. For spatial heterogeneity, GWR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.887   local coefficients vary by location   outcome: Y_value, predictors: X1, X2

>> COMMENTARY (narration):
    Assuming the relationship is NOT constant across space, we estimated SEPARATE coefficients for each location (R^2 =
    0.89). Unlike "global" regression, GWR fits a separate model at each point with local overlap/weights; this answers
    "how does this variable's effect vary by region?" and maps spatial heterogeneity. In the lab/environmental studies
    it is powerful where the relationship differs by region.

====================================================================================

#108  Nested Mixed Model
    file: 108_nested_lmm_R_P_F.xlsx
  >> SCENARIO (narration):
    We model a value in a nested design (upper>lower>block). For nested hierarchies,
    nested LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): lower_group_no
    - Factor (categorical): upper_group
    - Factor (categorical): block

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper_group (P) effect: F = 27.44, p < .001 ***   lower_group_no (R) effect: F = 6.11, p < .001 ***
    nested variance components (block within P)

>> COMMENTARY (narration):
    We modeled a value in a nested design (upper group > lower group > block): both the upper level (F = 27.44,
    p < .001) and the lower/replication level (F = 6.11, p < .001) make significant contributions. Nested LMM correctly
    handles hierarchies where sub-units are nested within super-units (each block belongs to only one group); by
    partitioning variance into levels it shows each layer's share. In the lab it is used to correctly separate effects
    in batch/sample/replicate hierarchies.

====================================================================================

#109  Crossed Mixed Model
    file: 109_crossed_lmm_A_B.xlsx
  >> SCENARIO (narration):
    We model a design where two random factors are crossed. For two independent
    classification axes, crossed LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): factor_a
    - 2nd Factor: factor_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    A x B interaction: F = 14.67, p < .001 ***   main effects ns (A: p=0.767, B: p=0.480)   outcome: value

>> COMMENTARY (narration):
    We modeled a design where two random factors are crossed rather than nested (each level of A pairs with each level
    of B): while the main effects are non-significant (A: p=0.77, B: p=0.48), the A x B interaction is very strong
    (F = 14.67, p < .001) -- so the effect depends on the COMBINATION of factors. Crossed LMM, unlike nested, handles
    two independent grouping axes (e.g. method x batch, each method in each batch) at once. In the lab it is the right
    choice for jointly analyzing the effects and interaction of two independent classification axes.

====================================================================================

#110  Kernel Density (KDE) Map
    file: 110_KDE_sample_density.xlsx
  >> SCENARIO (narration):
    We produce a continuous density surface from sample locations. For a
    sample-concentration map, KDE is appropriate.
  >> VARIABLE SELECTION:
    - Value: measurement_value
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 95 samples   measurement-weighted spatial density surface (continuous heat map)

>> COMMENTARY (narration):
    Using sample locations (n = 95) we produced a continuous density surface (a KDE heat map): the point distribution
    was turned into a smooth "density" surface. KDE answers "where is it dense?" from scattered point data as a
    continuous map; it shows the trend rather than individual points. In the lab/field studies it is the core spatial
    tool for visually mapping sample/measurement concentrations and identifying empty/dense zones.

====================================================================================

#111  Hexbin Density Map
    file: 111_Hexbin_lab_noktalari.xlsx
  >> SCENARIO (narration):
    We aggregate measurement points into hexagonal cells to show density. For the
    over-plotting problem, hexbin is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 150 points   spatial density via aggregation into hexagonal cells

>> COMMENTARY (narration):
    We colored measurement points (n = 150) by counting within hexagonal cells. Hexbin aggregates many overlapping
    points (the over-plotting problem) into a regular hexagonal grid to show density clearly; compared with squares it
    carries less directional bias. In the lab/field studies it is used to turn dense point clouds (sample locations)
    into a readable density map and to compare spatial concentrations.

====================================================================================

#112  Moran's I
    file: 112_Morans_I_measurement_autocorrelation.xlsx
  >> SCENARIO (narration):
    We test whether the measurement value is distributed randomly or in clusters
    across space. For spatial autocorrelation, Moran's I is appropriate.
  >> VARIABLE SELECTION:
    - Value: measurement_value
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.777   z = 13.83   p < .001 ***   (k-NN neighbors, k = 6)   variable: measurement_value

>> COMMENTARY (narration):
    We tested whether the measurement value is distributed randomly or in clusters across space with Moran's I: I =
    0.78, z = 13.83, p < .001 -- very strong positive spatial autocorrelation, i.e. similar measurement values cluster
    geographically. Moran's I quantifies "Tobler's first law" (near things are similar). In the lab/environmental
    studies it is the first test for detecting the geographic clustering of measurements and for deciding whether a
    spatial model is needed.

====================================================================================

#113  Getis-Ord Gi*
    file: 113_Getis_Ord_high_value_hotspot.xlsx
  >> SCENARIO (narration):
    We map locally where the measurement value clusters high/low. For hot/cold
    spots, Getis-Ord is appropriate.
  >> VARIABLE SELECTION:
    - Value: measurement_value
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Hot spots = 29   Cold spots = 28   n = 68   (k-NN k = 8, binary weights)   variable: measurement_value

>> COMMENTARY (narration):
    We mapped locally WHERE the measurement value clusters high/low with Getis-Ord Gi*: 29 significant hot spots
    (high-value clusters) and 28 cold spots (low-value clusters). While Moran's I states the overall clustering, Gi*
    shows its location -- for each point it tests "is the surrounding area high or low?". In the lab/environmental
    studies it is used to pinpoint high/low-value zones (hot/cold spots) for targeted sampling/monitoring.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_sample_kumeleri.xlsx
  >> SCENARIO (narration):
    We cluster sample locations with density-based DBSCAN. For spatial clusters and
    outlier locations, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Clusters = 5   Noise (outliers) = 9   n = 115   (location: lat, lon)

>> COMMENTARY (narration):
    We clustered sample locations (n = 115) with density-based DBSCAN: 5 natural geographic clusters and 9 "noise"
    (scattered, cluster-less) points. DBSCAN does not require the cluster count in advance, finds clusters of any
    shape, and separates sparse points as outliers -- ideal for spatial clustering. In the lab/field studies it is used
    to detect region-level natural concentrations (sample clusters) and to isolate isolated locations; it completes our
    descriptive spatial analysis series.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_reaction_yield_pct.xlsx
  >> SCENARIO (narration):
    We follow 40 samples measured at three temperature levels (low, medium, high); each belongs to one of two catalyst groups (catalyst_A / catalyst_B). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on reaction yield (%).
  >> VARIABLE SELECTION:
    - Dependent variable: reaction_yield_pct
    - Subject ID: sample_id
    - Between-subjects factor: catalyst
    - Within-subjects factor: temperature

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (catalyst): F(1,38) = 6.10  p = 0.018  np2 = 0.138
    Within (temperature):  F(2,76) = 218.68  p < .001  np2 = 0.852
    Interaction:         F(2,76) = 7.19  p = 0.001  np2 = 0.159
    Mauchly W = 0.962  p = 0.478   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a catalyst x temperature mixed design we analyzed reaction yield (%) for 40 samples (120 observations). The interaction is significant (F(2,76) = 7.19, p = 0.001, np2 = 0.159) ** -- the two groups' change across temperature differs in magnitude. The between-subjects main effect (catalyst_A vs catalyst_B) is F = 6.10, p = 0.018; the within-subjects main effect (low/medium/high) is F = 218.68, p < .001. Mauchly's test p = 0.478, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each temperature level. In physical sciences, the mixed design is the standard analysis for comparing catalysts across temperature conditions.

====================================================================================
