====================================================================================
MerQur - Beklioglu_Limnoloji - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Beklioglu_Limnoloji/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 001_descriptive_konya_basin_lake_inventory.xlsx
  >> SCENARIO (narration):
    In this study we compiled a baseline inventory of 120 lakes across the Konya
    Basin to understand the physical and trophic character of the region's
    standing waters. For each lake we recorded morphometric traits such as
    area_km2, max_depth_m, mean_depth_m and elevation_m, alongside key water-
    quality indicators including ph, conductivity_uS_cm, chla_ugL, secchi_m,
    TP_ugL and TN_mgL. Before any modelling, we simply want to summarize the
    central tendency, spread and range of these variables to know what a typical
    lake looks like and where the extremes lie. Descriptive Statistics is the
    natural first step here because it profiles each measured variable without
    assuming any relationship or group structure.
  >> VARIABLE SELECTION:
    - Category: lake_name
    - Variable: area_km2
    - Variable: max_depth_m
    - Variable: mean_depth_m
    - Variable: elevation_m
    - Variable: ph
    - Variable: conductivity_uS_cm
    - Variable: chla_ugL
    - Variable: secchi_m
    - Variable: TP_ugL
    - Variable: TN_mgL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 120 lakes (Konya Closed Basin inventory)
    area_km2   : mean = 21.18   min-max = 0.9 - 402.1      |   max_depth_m : mean = 9.51   (0.8 - 34.6)
    ph         : mean = 8.38    (7.1 - 10.0)               |   chla_ugL    : mean = 22.0   (1.5 - 223.5)
    secchi_m   : mean = 1.24    (0.1 - 4.4)                |   TP_ugL      : mean = 62.1   (3 - 297)
    TN_mgL     : mean = 1.46    (0.1 - 5.6)

>> COMMENTARY (narration):
    We first drew the overall limnological picture of the basin: 120 lakes ranging from 0.9 to 402 square
    kilometres -- from small ponds to large lakes. Mean maximum depth is 9.5 m, so most are shallow lakes. A mean
    pH of 8.4 is mildly alkaline -- expected in the carbonate, closed-basin lakes of Anatolia. The striking values
    are the trophic indicators: mean chlorophyll-a 22, total phosphorus 62 micrograms, Secchi depth only 1.2 m.
    These say most lakes in the basin are eutrophic -- heavily nutrient-loaded, with low water clarity. The very
    high chlorophyll and phosphorus maxima (223 and 297) show some lakes are under severe eutrophication. This
    descriptive table sets the stage for the analyses that follow -- especially phosphorus thresholds and restoration effects.

====================================================================================

#2  Normality Tests
    file: 002_normality_eymir_mogan_chla_20yil.xlsx
  >> SCENARIO (narration):
    We have a 20-year monitoring record from two connected Ankara lakes, Eymir
    and Mogan, and we want to know whether their chlorophyll-a values are
    suitable for parametric analysis. The dataset holds monthly chla_ugL
    together with TP_ugL, secchi_m and WL_water_level_m across many years.
    Because nutrient and algal data are often strongly right-skewed, we must
    check the distribution of chla_ugL and the related water-quality variables
    before choosing t-tests or ANOVA. Normality Tests fit this purpose exactly,
    telling us whether each variable departs from a normal distribution and
    whether transformation will be needed.
  >> VARIABLE SELECTION:
    - Category: lake
    - Variable: chla_ugL
    - Variable: TP_ugL
    - Variable: secchi_m
    - Variable: WL_water_level_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chla_ugL : Shapiro-Wilk = 0.777  p < .001   Not Normal
    TP_ugL   : Shapiro-Wilk = 0.848  p < .001   Not Normal
    secchi_m : Shapiro-Wilk = 0.997  p = 0.962   Normal

>> COMMENTARY (narration):
    We tested whether three variables in the 20-year Eymir and Mogan series are normally distributed. The result is
    a typical limnology pattern: chlorophyll-a and total phosphorus deviate significantly from normality (p near
    zero), while Secchi depth is almost perfectly normal (p = 0.96). This is no surprise -- chlorophyll and phosphorus
    are typically right-skewed in aquatic systems: most values low, but very high values at eutrophic peaks. The
    practical takeaway is clear: we can use parametric tests (t-test, ANOVA) on Secchi; but for chlorophyll and
    phosphorus, nonparametric methods (Mann-Whitney, Kruskal-Wallis) or a log-transform are more appropriate.
    This is a pre-step that should not be skipped when analysing water-quality series.

====================================================================================

#3  One-Sample t-Test
    file: 003_one_sample_t_beysehir_TP_WFD_threshold.xlsx
  >> SCENARIO (narration):
    Lake Beysehir is a major drinking-water reservoir, and we want to know
    whether its total phosphorus concentration meets the Water Framework
    Directive reference threshold. We have 36 seasonal measurements of TP_ugL
    recorded across spring, summer and autumn, and the regulatory threshold is
    stored in test_obtained_WFD_threshold. The question is whether the mean
    TP_ugL differs significantly from that single fixed target value. A One-
    Sample t-Test is appropriate because we are comparing the mean of one
    continuous variable against a known constant.
  >> VARIABLE SELECTION:
    - Test Variable: TP_ugL
    - Test Value: test_obtained_WFD_threshold

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(35) = 7.1108   p < .001 ***   Cohen d = 1.185 (Large)
    Mean TP = 68.14 µg/L   (test mu = 50, WFD threshold)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the total phosphorus level of Lake Beysehir against the 50-microgram reference threshold the Water
    Framework Directive (WFD) sets for good water quality. The result is worrying: mean TP is 68.1 micrograms,
    clearly above the threshold -- t(35) = 7.11, p below one in a thousand. The effect size is very large at Cohen
    d = 1.19. So the lake's phosphorus load exceeds the "good status" threshold by more than chance can explain. The
    practical meaning matters: by WFD criteria the lake is nutrient-degraded; basin-scale management (agricultural
    runoff control, wastewater treatment) is needed to cut phosphorus inputs. The one-sample t-test is the correct
    way to compare an indicator against a known management threshold -- a very common question in limnology.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent Samples t-Test
    file: 004_independent_t_warm_cold_mesocosm_chla.xlsx
  >> SCENARIO (narration):
    In a mesocosm experiment we simulated two climate scenarios, warm and cold,
    to test how warming affects algal growth in shallow lakes. Across 72 tanks
    we measured chla_ugL together with the realized temperature_C in each tank.
    The core research question is whether mean chlorophyll-a differs between the
    warm and cold climate_region groups. An Independent Samples t-Test is the
    right choice because we are comparing one continuous outcome between two
    independent groups.
  >> VARIABLE SELECTION:
    - Group: climate_region
    - Test Variable: chla_ugL
    - Test Variable: temperature_C

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(70) = 6.618   p < .001 ***   Cohen d = 1.560 (Large)
    warm: n=36, mean chla = 40.62   |   cold: n=36, mean chla = 24.15   |   H0 REJECTED

>> COMMENTARY (narration):
    We tested the effect of climate warming on lake productivity by comparing warm- and cold-climate mesocosm tanks.
    The result is striking: mean chlorophyll-a is 40.6 in warm tanks versus 24.1 micrograms in cold ones -- nearly
    double under warm conditions. The difference is highly significant (t(70) = 6.62, p < .001) and the effect size
    is large (d = 1.56). This experimentally confirms a core prediction of climate-change limnology: warming
    accelerates algal production and eutrophication, even at the same nutrient level. The independent-samples t-test
    is the standard way to compare a continuous measure between two independent groups -- here, two climate treatments.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired t-Test
    file: 005_paired_t_eymir_restoration_TP.xlsx
  >> SCENARIO (narration):
    Lake Eymir underwent a restoration effort, and we want to evaluate whether
    sewage diversion lowered phosphorus loading at fixed monitoring stations.
    For each of 30 stations we have TP_before_1994 and TP_post_2002, measured at
    the same location before and after the intervention. The question is whether
    total phosphorus changed significantly within the same stations over time. A
    Paired t-Test fits because the two measurements are dependent observations
    taken from the same stations.
  >> VARIABLE SELECTION:
    - Before: TP_before_1994
    - After: TP_post_2002

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Paired t(29) = 20.3766   p < .001 ***   Cohen d_z = 3.720 (Large)
    Mean difference (pre-1994 − post-2002 TP) = 70.04 µg/L   H0 REJECTED

>> COMMENTARY (narration):
    We evaluated the success of Lake Eymir's restoration by comparing total phosphorus at the same stations before
    (pre-1994) and after (post-2002) the intervention. The result is dramatic: mean TP dropped by 70 micrograms,
    t(29) = 20.38, p < .001, with an enormous effect size (d_z = 3.72). Because we measured the same stations twice,
    the paired t-test is correct -- it removes between-station variability and focuses on the within-station change.
    This is concrete evidence that the restoration (sewage diversion, biomanipulation) worked. The paired t-test is
    the right tool for before/after monitoring at the same sites -- one of the most common designs in applied limnology.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 006_one_way_anova_salinity_zooplankton.xlsx
  >> SCENARIO (narration):
    Salinity strongly shapes lake food webs, so we examined how zooplankton
    biomass responds across a gradient of salinity classes. Our 112 samples are
    grouped into fresh, oligo_to, meso_to and Poli_to salinity_class, with the
    measured salinity_psu and zooplankton_biomass_ugDW_L recorded for each. We
    want to know whether mean zooplankton_biomass_ugDW_L differs among these
    four salinity groups. A One-Way ANOVA is appropriate because we are
    comparing a continuous response across more than two independent groups.
  >> VARIABLE SELECTION:
    - Factor: salinity_class
    - Dependent: zooplankton_biomass_ugDW_L

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 108) = 135.7531   p < .001 ***   η² = 0.7904
    DECISION: H0 REJECTED (4 salinity classes: fresh / oligo / meso / poly)

>> COMMENTARY (narration):
    We compared chlorophyll across four salinity classes (fresh, oligo-, meso-, polyhaline) with one-way ANOVA. The
    result is overwhelming: F(3,108) = 135.75, p < .001, and an effect size of η² = 0.79 -- salinity class explains
    almost 80% of the variance in chlorophyll. So salinity is a powerful structuring factor for algal biomass in
    these lakes. ANOVA tests whether at least one group mean differs; the huge effect here calls for a post-hoc test
    to see which classes differ. Salinity gradients -- driven by evaporation and irrigation in closed basins --
    profoundly reshape primary production, and ANOVA quantifies that effect clearly.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 007_two_way_anova_nutrient_suSeviyesi_mesocosm.xlsx
  >> SCENARIO (narration):
    In a factorial mesocosm experiment we manipulated both nutrient loading and
    water level to see how they jointly drive algal blooms. Across 90 tanks,
    chla_ugL was measured under two nutrient_loading levels, low_N and high_N,
    crossed with three water_level conditions, low_WL, medium_WL and high_WL. We
    want to test the main effects of nutrient loading and water level, and
    crucially whether the two factors interact in shaping chla_ugL. A Two-Way
    ANOVA fits because there are two categorical factors and one continuous
    outcome.
  >> VARIABLE SELECTION:
    - Factor: nutrient_loading
    - Factor: water_level
    - Dependent: chla_ugL

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    nutrient_loading (main) : F(1) = 332.88   p < .001 ***   η²p = 0.80
    water_level (main)      : F(2) =  42.70   p < .001 ***   η²p = 0.50
    nutrient x water_level  : F(2) =  33.41   p < .001 ***   (interaction SIGNIFICANT)

>> COMMENTARY (narration):
    We examined the joint effect of two factors -- nutrient loading and water level -- on chlorophyll with two-way
    ANOVA. Both main effects are highly significant (nutrient F = 332.88, η²p = 0.80; water level F = 42.70), but
    the key finding is the significant interaction (F = 33.41, p < .001): the effect of nutrients depends on water
    level. In other words, high nutrient loading hits chlorophyll hardest under low water levels -- a critical
    insight for shallow-lake management under drought. Two-way ANOVA's strength is exactly this: it reveals not just
    independent effects but how two factors combine. The significant interaction means we cannot interpret either
    factor in isolation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated Measures ANOVA
    file: 008_RM_anova_mesocosm_8wk_chla.xlsx
  >> SCENARIO (narration):
    We ran an 8-week warming mesocosm trial to track how chlorophyll-a develops
    over time under heating. Twelve tanks were assigned to either a control or a
    heat_increase treatment, and chla was recorded weekly from week_1_chla
    through week_8_chla. The question is whether chlorophyll-a changes
    significantly across the eight weeks and whether that time trajectory
    differs between treatments. A Repeated Measures ANOVA is suited here because
    the same tanks are measured repeatedly over time.
  >> VARIABLE SELECTION:
    - Between Factor: treatment
    - Within (Week 1): week_1_chla
    - Within (Week 2): week_2_chla
    - Within (Week 3): week_3_chla
    - Within (Week 4): week_4_chla
    - Within (Week 5): week_5_chla
    - Within (Week 6): week_6_chla
    - Within (Week 7): week_7_chla
    - Within (Week 8): week_8_chla

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(7, 77) = 12.1575   p < .001 ***   η²p = 0.5250   (n = 12 tanks)
    DECISION: H0 REJECTED

>> COMMENTARY (narration):
    We tracked chlorophyll in the same 12 tanks across 8 months and tested for change with repeated-measures ANOVA.
    Because the same tanks are measured repeatedly, observations are dependent; RM-ANOVA accounts for this within-
    subject correlation, which an ordinary ANOVA would ignore. The result is significant (F(7,77) = 12.16, p < .001,
    η²p = 0.53): chlorophyll changes substantially over the months -- a strong seasonal/temporal dynamic. RM-ANOVA
    is the correct design for longitudinal experiments where the same units are followed through time, and is far
    more powerful than treating each time point as independent.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 009_manova_mesocosm_4yanit.xlsx
  >> SCENARIO (narration):
    In a mesocosm experiment crossing nutrient addition and warming, we want to
    know how the treatments jointly affect the lake ecosystem across several
    correlated response variables at once. Across 100 tanks assigned to control,
    nutrient, heat and nutrient_heat treatments, we measured chla_ugL, TP_ugL,
    TN_mgL and DO_mgL together. Because these four responses are biologically
    linked, we ask whether the treatment groups differ on the combined set of
    outcomes rather than testing each separately. A MANOVA fits because it
    evaluates group differences across multiple dependent variables
    simultaneously.
  >> VARIABLE SELECTION:
    - Factor: treatment
    - Dependent: chla_ugL
    - Dependent: TP_ugL
    - Dependent: TN_mgL
    - Dependent: DO_mgL

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda -> F = 57.7335   p < .001 ***   (n = 100)
    Dependent: chla_ugL, TP_ugL, TN_mgL, DO_mgL   |   Factor: treatment

>> COMMENTARY (narration):
    Here we have four correlated response variables -- chlorophyll, phosphorus, nitrogen and dissolved oxygen -- and
    one treatment factor. Running four separate ANOVAs would inflate the error rate and ignore the correlations
    among responses. MANOVA tests all four jointly: Wilks' Lambda gives F = 57.73, p < .001, so the treatment shifts
    the multivariate response profile strongly. MANOVA is the right approach when an intervention affects several
    related outcomes at once -- it controls the overall error rate and detects effects spread across variables that
    univariate tests might miss. After a significant MANOVA, follow-up univariate tests show which variables drive it.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 010_ancova_saltmarsh_blue_carbon_depth_kov.xlsx
  >> SCENARIO (narration):
    In coastal salt marshes we want to compare organic carbon storage among
    three vegetation types while accounting for sampling depth. Across 90 cores
    from Salicornia, Juncus and Phragmites ecosystem_type stands, we recorded
    OC_percent along with depth_cm and dry_density_gcm3, and the marshes span
    the aegean and mediterranean coastal strips. Since deeper samples naturally
    carry different carbon, we want to compare mean OC_percent among vegetation
    types after adjusting for depth_cm as a covariate. An ANCOVA fits because it
    tests group differences in a continuous outcome while controlling for a
    continuous covariate.
  >> VARIABLE SELECTION:
    - Factor: ecosystem_type
    - Dependent: OC_percent
    - Covariate: depth_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ecosystem_type (factor): F(2) = 171.80   p < .001 ***   η²p = 0.800 (large)
    Covariate: depth_cm (adjusted for depth)   |   DV: OC_percent (organic carbon %)   H0 REJECTED

>> COMMENTARY (narration):
    We compared sediment organic carbon across ecosystem types while controlling for sediment depth -- a potential
    confounder, since deeper sediments differ in carbon. ANCOVA statistically holds depth constant. The result: even
    after adjusting for depth, ecosystem types differ significantly in organic carbon (F(2) = 171.80, p < .001),
    with a very large effect (η²p = 0.80). So the difference we observe is not a by-product of depth differences; it
    is a genuine ecosystem-type effect. Controlling confounders like depth is essential in wetland and sediment
    studies -- ANCOVA does this fairly and is the correct tool when a continuous covariate must be accounted for.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 011_bootstrap_CI_small_N_lake_chla.xlsx
  >> SCENARIO (narration):
    How do we report a reliable average chlorophyll-a concentration when our
    sampling effort across these lakes is genuinely small? Across just 22 sites
    spanning lakes such as Cavuscu, Sugla, Hotamis, mogan, Burdur, and eymir, we
    measured chla_ugL and want a trustworthy confidence interval for the mean.
    Because the sample size is small and chla_ugL may not be normally
    distributed, classical normal-theory intervals are shaky. A Bootstrap
    Confidence Interval resamples our observed chla_ugL values many times,
    building an empirical interval that does not lean on distributional
    assumptions. This lets us state the plausible range of mean chlorophyll-a
    with honesty about our limited N.
  >> VARIABLE SELECTION:
    - Measurement variable: chla_ugL
    - Grouping label: lake_name

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed median chlorophyll-a = 16.08 ug/L
    95% Bootstrap CI = (9.18 , 24.55)   |   small sample (low N)

>> COMMENTARY (narration):
    Sometimes we have very few lakes/samples and a skewed distribution -- classic formulas become unreliable. Here,
    for a small chlorophyll-a sample, we derived the confidence interval of the median by bootstrapping: the observed
    median is 16.1, but the plausible range for the true median is 9.2 to 24.5. This width tells us two things:
    uncertainty is high because the sample is small, and the chlorophyll distribution is skewed. Bootstrap estimates
    the interval without distributional assumptions, through thousands of resamples -- a reliable tool for limnologists
    working with little data.

====================================================================================

#12  Permutation Test
    file: 012_permutasyon_zoo_dimension_temperature.xlsx
  >> SCENARIO (narration):
    Does water temperature shape the body size of zooplankton in our lakes? We
    compared zoo_dimension_um between two thermal regimes, a cold_15C group and
    a warm_24C group, across 70 individuals to test the classic temperature-size
    rule. Because we are contrasting a single continuous measurement between two
    independent temperature groups and do not want to assume normality of the
    size distribution, a Permutation Test is ideal. By repeatedly shuffling the
    temperature_group labels and recomputing the difference in mean
    zoo_dimension_um, we build an exact null distribution and judge whether the
    observed thermal difference is real or just chance.
  >> VARIABLE SELECTION:
    - Dependent variable: zoo_dimension_um
    - Group: temperature_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Significant mean difference between the two (temperature) groups (p < 0.05).
    Zooplankton body size differs by temperature group.   H0 REJECTED

>> COMMENTARY (narration):
    We tested the effect of temperature on zooplankton body size with a permutation test. Permutation shuffles the
    group labels thousands of times to measure whether the observed difference could arise by chance -- requiring no
    distributional assumption. The result is significant: body sizes under warm versus cold conditions differ more
    than chance allows. This matches the ecological "temperature-size rule": organisms tend to be smaller in warmer
    waters. For small samples and non-normal distributions, the permutation test is a robust alternative to the t-test.

====================================================================================

#13  Multiple Comparison Corrections
    file: 013_multiple_comparison_5havza_chla_tukey.xlsx
  >> SCENARIO (narration):
    When we compare chlorophyll-a across five major basins, how do we avoid
    being fooled by all those pairwise tests? We measured chla_ugL in 110 lake
    samples drawn from the Konya, large_Menderes, north_aegean, Marmara, and
    east_black_sea basins and want to know which basins truly differ in trophic
    richness. Running every basin-to-basin comparison inflates the false-
    positive rate, so we apply Multiple Comparison Corrections such as Tukey
    adjustment. This controls the family-wise error while still pinpointing
    which basin pairs show genuinely different chla_ugL levels.
  >> VARIABLE SELECTION:
    - Dependent variable: chla_ugL
    - Group: basin

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 basins vs reference: all p-values stayed significant under every method (Bonferroni / Holm / FDR ...):
    TOTAL SIGNIFICANT = 4/4 (in every correction column)

>> COMMENTARY (narration):
    When we compare the chlorophyll levels of five basins pairwise, we run many tests -- and each carries a false-
    positive risk. Multiple-comparison correction controls that risk. Here all four comparisons against the reference
    basin stayed significant under every correction method, from Bonferroni to FDR (4/4). This is a strong result:
    the chlorophyll differences among basins are so clear that even the most conservative correction does not
    eliminate them. In multi-lake comparisons, skipping this step risks reporting spurious differences.

====================================================================================

#14  Mann-Whitney U Test
    file: 014_mann_whitney_phytoplankton_density.xlsx
  >> SCENARIO (narration):
    Does salinity influence how densely phytoplankton populate our waters? We
    compared phytoplankton_cell_L between low and high salt_level conditions
    across 64 samples to see whether saltier lakes support different algal
    abundance. Phytoplankton densities are typically right-skewed with extreme
    blooms, so a test that assumes normality would be misleading. The Mann-
    Whitney U Test compares the rank distributions of phytoplankton_cell_L
    between the two independent salt_level groups, giving a robust verdict
    without distributional assumptions.
  >> VARIABLE SELECTION:
    - Dependent variable: phytoplankton_cell_L
    - Group: salt_level

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Mann-Whitney U = 172.00   p < .001 ***   r = .66
    Phytoplankton density differs significantly by salinity level.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether phytoplankton cell density differs by salinity level with the nonparametric Mann-Whitney test
    -- because density data are heavily skewed. The result is highly significant (U = 172, p < .001) with a strong
    effect size (r = .66). So phytoplankton density clearly differs between low- and high-salinity lakes. This is
    ecologically important: salinization (especially evaporation- and irrigation-driven in closed basins) reshapes
    plankton communities. For skewed density data, Mann-Whitney is more reliable than the t-test.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank Test
    file: 015_wilcoxon_conductivity_before_post.xlsx
  >> SCENARIO (narration):
    Did an intervention actually change the conductivity of our lakes? For 24
    lakes we recorded conductivity_before_mS_cm and conductivity_post_mS_cm on
    the very same waterbodies, giving us paired before-and-after measurements.
    Because the two readings come from the same lakes and conductivity changes
    may not follow a normal distribution, a paired nonparametric test is the
    right choice. The Wilcoxon Signed-Rank Test evaluates the signed differences
    between conductivity_before_mS_cm and conductivity_post_mS_cm to detect a
    consistent shift.
  >> VARIABLE SELECTION:
    - Before: conductivity_before_mS_cm
    - After: conductivity_post_mS_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilcoxon W = 15.00   p < .001 ***   r = 0.88   n (nonzero) = 24
    Conductivity changed significantly before vs after.   H0 REJECTED

>> COMMENTARY (narration):
    We compared electrical conductivity measured before and after an intervention/period at the same stations with
    the paired Wilcoxon test. The result is very strong: W = 15, p < .001, effect size r = 0.88. Conductivity
    changed in the same direction and by a large amount at nearly all stations. Conductivity is a direct proxy for
    the ionic load entering the water (salinization, pollution); such a consistent shift signals a real
    hydrochemical change in the system. For paired measures with skewed distributions, Wilcoxon is the right choice.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 016_kruskal_wallis_bird_diversity_term.xlsx
  >> SCENARIO (narration):
    Has bird community richness around our lakes changed over the decades? We
    grouped 112 observations of bird_type_richness into four monitoring terms,
    2003-2007, 2008-2012, 2013-2017, and 2018-2022, to track ecological change
    through time. Richness counts are bounded and often non-normal, so a rank-
    based test across more than two groups is appropriate. The Kruskal-Wallis
    Test compares the distribution of bird_type_richness across the four term
    groups to reveal whether biodiversity differs among periods.
  >> VARIABLE SELECTION:
    - Dependent variable: bird_type_richness
    - Group: term

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(3) = 58.84   p < .001 ***   eta-squared_H = 0.52
    Bird species richness differs significantly by season/term.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether wetland bird species richness varies across periods/seasons with Kruskal-Wallis -- the
    nonparametric ANOVA for more than two groups. The result is highly significant (H(3) = 58.84, p < .001) with a
    very large effect (eta-squared = 0.52): more than half the variance is explained by period. Wetlands host very
    different bird communities during migration, breeding and wintering; this result quantifies that seasonal
    dynamic. Because count data are not normally distributed, Kruskal-Wallis is more appropriate than classic ANOVA.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 017_friedman_mevsimler_chla_aynilgol.xlsx
  >> SCENARIO (narration):
    Does chlorophyll-a follow a seasonal rhythm within the same lakes? Across 18
    lakes we measured spring_chla, summer_chla, autumn_chla, and winter_chla,
    giving four repeated chla readings per waterbody. Since the same lakes are
    measured in every season and chlorophyll values are skewed, a nonparametric
    repeated-measures test fits best. The Friedman Test ranks the four seasonal
    chla measurements within each lake to test for a consistent seasonal pattern
    in algal productivity.
  >> VARIABLE SELECTION:
    - Repeated measure: spring_chla
    - Repeated measure: summer_chla
    - Repeated measure: autumn_chla
    - Repeated measure: winter_chla

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Friedman chi-square(3) = 35.87   p < .001 ***   Kendall W = 0.66   n = 18
    The 4-season chlorophyll of the same lakes differs significantly.   H0 REJECTED

>> COMMENTARY (narration):
    We compared chlorophyll-a measured across four seasons in the same lakes with the Friedman test -- the
    nonparametric counterpart of repeated-measures ANOVA. The result is highly significant (chi-square(3) = 35.87,
    p < .001), and Kendall W = 0.66 indicates high consistency of the seasonal ranking. So lakes follow a consistent
    seasonal chlorophyll pattern: typically a summer peak and winter low. Because we measure the same lakes
    repeatedly, the observations are dependent; Friedman accounts for that correctly. It is the right test for
    seasonal monitoring data.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 018_binomial_clear_status_ratio.xlsx
  >> SCENARIO (narration):
    What fraction of our lake-years can be classified as clear-water, and is
    that proportion different from an even split? Across 80 lake-year records we
    coded clear_status_1_0 as one for clear and zero for turbid, alongside
    TP_ugL as context. To judge whether the share of clear lakes departs from a
    hypothesized 50 percent, we treat each record as a binary outcome. The
    Binomial Test compares the observed count of clear_status_1_0 successes
    against the expected probability, telling us if clear conditions are over-
    or under-represented.
  >> VARIABLE SELECTION:
    - Binary outcome: clear_status_1_0
    - Covariate: TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed proportion (clear) = 0.6625   test proportion = 0.50
    p = 0.0049 **   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the proportion of lakes in a "clear water" state equals 50% (chance). The observed proportion
    is 0.66 -- two-thirds of lakes are clear. This differs significantly from 50% (p = 0.005). Shallow lakes have two
    alternative stable states: "clear" (macrophyte-dominated) and "turbid" (algae-dominated). This result shows the
    clear regime dominates among the monitored lakes -- a positive sign for restoration and macrophyte protection.
    The binomial test is the simplest way to compare a single binary proportion against a known reference.

====================================================================================

#19  Sign Test
    file: 019_sign_test_water_level_change.xlsx
  >> SCENARIO (narration):
    Have our lakes risen or fallen in water level over fifteen years? For 30
    lakes we recorded WL_2005_m and WL_2020_m, the water level at each lake in
    2005 and again in 2020. We only trust the direction of change, not its exact
    magnitude, because measurement scales vary and the differences are far from
    normal. The Sign Test simply counts how many lakes went up versus down
    between WL_2005_m and WL_2020_m to test for a systematic level change.
  >> VARIABLE SELECTION:
    - Before: WL_2005_m
    - After: WL_2020_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive (2005 > 2020) = 21   Negative (2005 < 2020) = 9
    p = 0.0428 *   H0 REJECTED  (2005 levels significantly higher)

>> COMMENTARY (narration):
    We compared the 2005 and 2020 water levels of the same lakes with the paired sign test. In 21 of 30 lakes the
    water level was higher in 2005, and rose in only 9 (p = 0.043). So over 15 years water level fell in the majority
    of lakes. This statistically confirms one of the most critical problems of Anatolia's closed basins -- over-
    abstraction and climate-driven drying. The sign test looks only at the direction of change (up/down), not its
    magnitude; it is a robust, assumption-free test -- a coarse but sturdy trend indicator.

====================================================================================

#20  Runs Test
    file: 020_runs_test_chla_monthly_rastgelelik.xlsx
  >> SCENARIO (narration):
    Is the month-to-month pattern of chlorophyll-a random, or does it cluster
    into runs of high and low values? Over 120 monthly observations ordered by
    month_no, we recorded chla_ugL to check whether algal blooms appear in non-
    random sequences. If high and low chla_ugL values clump together rather than
    alternating by chance, that signals temporal structure or seasonality. The
    Runs Test examines the sequence of chla_ugL relative to its median to test
    the randomness of its ordering over time.
  >> VARIABLE SELECTION:
    - Sequence variable: chla_ugL
    - Order: month_no

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 70   Expected E[R] = 61.0   Z = 1.65   p ~ 0.099
    H0 NOT REJECTED  (series accepted as random)

>> COMMENTARY (narration):
    We tested whether the monthly chlorophyll series is randomly distributed -- that is, whether successive values
    cluster or show a pattern -- with the Wald-Wolfowitz runs test. The 70 observed runs are close to the 61
    expected; Z = 1.65, p about 0.10. H0 is not rejected: the series is accepted as random, with no systematic
    clustering. This matters methodologically: if residuals/measurements are non-random (autocorrelated), many
    classic tests become invalid. Because randomness holds here, standard methods can be used safely. The runs test
    is ideal for this pre-check.

====================================================================================

#21  Chi-Square Test of Independence
    file: 021_chisquare_independence_macrophyte_turbidity.xlsx
  >> SCENARIO (narration):
    In this study we surveyed 220 lakes to ask whether the presence of submerged
    macrophytes is linked to water turbidity. For each lake we recorded its
    turbidity_class as clear, turbid, or medium, and noted macrophyte_present as
    a yes or no indicator. Because both variables are categorical, a Chi-Square
    Test of Independence lets us test whether macrophyte occurrence is
    statistically associated with turbidity category rather than being
    independent of it. This matters ecologically, since clear-water lakes are
    expected to support more rooted vegetation than turbid ones.
  >> VARIABLE SELECTION:
    - Row variable: turbidity_class
    - Column variable: macrophyte_present

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(2) = 65.91   p < .001 ***   Cramer V = 0.55 (Strong)
    Turbidity class is associated with macrophyte presence.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether water turbidity class is associated with macrophyte (aquatic plant) presence using chi-square.
    The result is very strong: chi-square(2) = 65.91, p < .001, Cramer V = 0.55. So turbidity and macrophytes are
    not independent -- clear lakes have macrophytes, turbid lakes do not. This is the core mechanism of shallow-lake
    ecology: macrophytes clarify the water (trapping sediment, competing for nutrients), and clear water in turn
    feeds the macrophytes -- a positive feedback. For association between two categorical variables, the chi-square
    test of independence is the standard method.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 022_chisquare_goodnessfit_goodness_zooplankton_turoran.xlsx
  >> SCENARIO (narration):
    Here we examined the relative abundance of five zooplankton taxa in a lake
    community to see whether they occur in the proportions predicted by theory.
    For each type, ranging from daphnia_pulex and daphnia_magna to
    Bosmina_longirostris, Cyclops_strenuus, and Diaptomus_castor, we tallied the
    observed_count and specified an expected_ratio. A Chi-Square Goodness-of-Fit
    Test compares the observed counts against the expected distribution to judge
    whether the community composition departs from the hypothesized pattern.
    This helps us decide if certain taxa are over- or under-represented in the
    sampled lake.
  >> VARIABLE SELECTION:
    - Category: type
    - Observed: observed_count
    - Expected proportion: expected_ratio

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(4) = 67.80   p < .001 ***   n = 800   k = 5 species
    Expected ratios: 0.30, 0.20, 0.25, 0.15, 0.10   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the species composition of the zooplankton community fits the expected (reference) proportions
    with the chi-square goodness-of-fit test. The observed distribution of 800 individuals deviates significantly
    from the expected ratios (chi-square(4) = 67.80, p < .001). So the community structure has drifted away from the
    "expected" balance -- some species more abundant than expected, others less. Such shifts usually stem from
    eutrophication, fish predation or temperature. The goodness-of-fit test is the right way to compare an observed
    category distribution against a theoretical expectation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 023_fisher_exact_rare_fish_microhabitat.xlsx
  >> SCENARIO (narration):
    In this small survey of 24 lake locations we asked whether a rare fish
    species prefers macrophyte beds over open water. Each record notes the
    microhabitat, either macrophyte_bed or open_water, alongside
    rare_fish_present as a presence or absence flag. Because the sample is small
    and expected cell counts are low, Fisher's Exact Test is the appropriate way
    to evaluate the association between microhabitat and rare fish presence. The
    result tells us whether rare fish occurrence is disproportionately tied to
    vegetated refuge habitat.
  >> VARIABLE SELECTION:
    - Group: microhabitat
    - Outcome: rare_fish_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher Exact p = 0.0361 *   Odds Ratio = 0.10  (95% CI: 0.014 - 0.693)   Phi = 0.51
    Rare-fish presence is associated with microhabitat.   H0 REJECTED

>> COMMENTARY (narration):
    We tested the association between the presence of a rare fish species and microhabitat type with Fisher's exact
    test -- because chi-square is unreliable for tables with few observations (small cells). The result is
    significant (p = 0.036), odds ratio 0.10: the chance of finding the rare fish in a given microhabitat is 10 times
    lower than in the other. Phi = 0.51 indicates a strong association. Rare/threatened species are usually tied to
    very specific habitats; this result quantifies that dependence. For small-sample conservation data, Fisher's
    exact test is the correct choice.

====================================================================================

#24  McNemar Test
    file: 024_mcnemar_macrophyte_before_post.xlsx
  >> SCENARIO (narration):
    We revisited 60 lake monitoring stations to evaluate whether macrophyte
    cover changed between two survey years. At each station we recorded
    macrophyte_2010 and macrophyte_2024 as paired presence or absence
    observations of the same site. Since the two measurements are paired on the
    same stations, McNemar's Test focuses on the discordant cases to test
    whether there was a significant shift in macrophyte occurrence over time.
    This reveals whether the lakes experienced net colonization or loss of
    submerged vegetation across the fourteen-year span.
  >> VARIABLE SELECTION:
    - Before: macrophyte_2010
    - After: macrophyte_2024

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar chi-square = 16.96   p < .001 ***   discordant: b = 24 (loss), c = 2 (gain)
    Cohen g = 0.85   H0 REJECTED

>> COMMENTARY (narration):
    We compared macrophyte presence in the same lakes in 2010 and 2024 -- paired binary data call for McNemar's test.
    The result is striking: macrophytes were lost in 24 lakes and gained in only 2 (chi-square = 16.96, p < .001). So
    over 14 years macrophyte cover declined dramatically. This is the classic signature of a shallow-lake "clear-to-
    turbid" regime shift -- macrophytes collapse under eutrophication and falling water levels. McNemar looks only at
    the cells that changed (the discordant pairs); it is the right way to measure before/after categorical change.

====================================================================================

#25  Cohen's Kappa
    file: 025_cohen_kappa_cladocera_uzmanlar_between.xlsx
  >> SCENARIO (narration):
    To check the reliability of taxonomic identification, two specialists
    independently classified 80 cladoceran samples. Each sample carries
    expert1_diagnosis and expert2_diagnosis, with each rater assigning one of
    four taxa: daphnia, Alona, Chydorus, or Bosmina. Cohen's Kappa measures the
    agreement between the two experts beyond what would be expected by chance,
    which is essential before pooling their identifications. A high kappa would
    confirm that the cladoceran classifications are consistent and trustworthy.
  >> VARIABLE SELECTION:
    - Rater 1: expert1_diagnosis
    - Rater 2: expert2_diagnosis

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's Kappa = 0.797 (substantial/strong agreement)   p_o = 0.85   p_e = 0.26
    z = 14.75   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We measured the agreement between two experts independently identifying Cladocera (water flea) species with
    Cohen's Kappa. Kappa = 0.80 is in the "substantial/strong" range -- observed agreement is 85%, but the chance-
    corrected Kappa is the better measure because it removes accidental agreement. In taxonomy this matters greatly:
    if species identifications are inconsistent between observers, all biodiversity data become suspect. Kappa = 0.80
    shows the identifications are reliable and reproducible. Reporting inter-rater reliability is the standard quality-
    assurance step in taxonomic studies.

====================================================================================

#26  Cochran-Mantel-Haenszel Test
    file: 026_cmh_macrophyte_group_salinity_strata.xlsx
  >> SCENARIO (narration):
    In this experiment we tested whether a treatment affects macrophyte
    establishment while accounting for differing salinity conditions across 240
    samples. Each record reports the group as control or treatment and the
    outcome result_macrophyte_present, stratified by layer_salinity into
    high_salt and low_salt strata. The Cochran-Mantel-Haenszel Test evaluates
    the treatment-outcome association while controlling for the salinity strata,
    so we are not misled by salinity as a confounder. This shows whether the
    treatment effect on macrophyte presence holds consistently across saline and
    fresher conditions.
  >> VARIABLE SELECTION:
    - Group: group
    - Outcome: result_macrophyte_present
    - Stratum: layer_salinity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi-square(1) = 32.44   p < .001 ***   Common OR (MH) = 5.06  (95% CI: 2.84 - 9.02)
    Breslow-Day chi-square(1) = 0.70, p = 0.40 (homogeneous OR)   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between group and macrophyte presence while controlling for salinity strata using the
    CMH test. Even holding salinity constant, the association is strong: common odds ratio 5.06 (p < .001) -- stripped
    of the salinity confounder, the odds of macrophyte presence in one group are 5 times those in the other. The
    Breslow-Day test is non-significant (p = 0.40), meaning this association is consistent across all salinity strata
    (homogeneous OR). CMH is the classic way -- very valuable in limnology -- to control for a third variable by
    stratification.

====================================================================================

#27  Log-Linear Analysis
    file: 027_log_lineer_treatment_term_result.xlsx
  >> SCENARIO (narration):
    Here we model how three categorical factors jointly shape water clarity
    outcomes in a controlled lake mesocosm study. The cross-classified frequency
    counts span treatment (control or heat), term (summer or winter), and result
    (turbid or clear). A Log-Linear Analysis fits the cell frequencies to reveal
    which main effects and interactions among treatment, season, and clarity
    best explain the observed counts. This lets us see, for example, whether
    heat treatment drives turbidity especially in the summer term.
  >> VARIABLE SELECTION:
    - Factor: treatment
    - Factor: term
    - Factor: result
    - Frequency: frequency

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Best model AIC = 48.77   Pearson chi-square = 0.47 (model fits the data very well)
    Factors: treatment (control/heat) x term (summer/winter) x result (turbid/clear)

>> COMMENTARY (narration):
    We analysed the contingency structure formed jointly by three categorical variables -- treatment (control/
    heating), term (summer/winter) and result (turbid/clear) -- with a log-linear model. The selected model fits
    perfectly (Pearson chi-square = 0.47, very low), with AIC = 48.77 as the most parsimonious model. Log-linear
    analysis lets us untangle interactions among more than two categorical variables (which variables are dependent,
    which interaction is significant) -- it is like ANOVA for categorical count data. It is powerful for summarising
    the multi-way outcomes of experimental warming studies.

====================================================================================

#28  Crosstabulation
    file: 028_cross_tablo_fish_lake_status.xlsx
  >> SCENARIO (narration):
    In this survey of 180 fish samples we explored the relationship between fish
    species and the ecological state of their lakes. Each sample records the
    fish_type, one of Sudak, Sazan, Turna, or Catalca, together with the
    lake_status as turbid or clear. A Crosstabulation summarizes how the species
    are distributed across turbid and clear lakes, giving a clear picture of
    which fish dominate under each water condition. This descriptive breakdown
    sets the stage for understanding species habitat preferences.
  >> VARIABLE SELECTION:
    - Row variable: fish_type
    - Column variable: lake_status

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 5.65   p = 0.130   Cramer V = 0.18 (Weak-moderate)
    NO significant association between fish type and lake status.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We examined the association between fish type and lake status with a crosstab and chi-square. This time the
    result is non-significant (chi-square(3) = 5.65, p = 0.13): there is not enough evidence that the fish-type
    distribution differs by lake status. Non-significant results are valuable too -- always expecting an association
    is a mistake. Here fish composition is not an indicator that distinguishes lake status; so the lake status must
    be driven by other factors (nutrients, macrophytes, depth). The crosstab is the most basic and readable way to
    see the co-occurrence of two categorical variables.

====================================================================================

#29  MR Frequency
    file: 029_MR_frequency_pondscape_function_choice.xlsx
  >> SCENARIO (narration):
    We surveyed 120 participants from several countries about the ecological
    functions they value in pondscapes, where each person could select multiple
    options. The data record each participant's country, drawn from turkey,
    sweden, germany, spain, and Cek_Cum, their age, and their multiple-response
    pondscape_fonksiyonlari_choice selections. A Multiple Response Frequency
    analysis tallies how often each pond function was chosen across all
    respondents, accounting for the fact that people picked more than one. This
    summarizes which pondscape functions are most widely appreciated by the
    public.
  >> VARIABLE SELECTION:
    - Multiple response set: pondscape_fonksiyonlari_choice

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 120   Respondents = 120 (100%)
    Pond/wetland functions (water provision, climate regulation, aesthetics, recreation...) selected as multi-choice

>> COMMENTARY (narration):
    Survey respondents were asked which functions of the pond/wetland landscape (pondscape) they value -- a
    "multiple response" question where several options can be ticked. All 120 respondents selected at least one
    function, so the response rate is 100%. Multiple-response frequency analysis shows how many times and by what
    percentage of respondents each function was chosen; the totals exceed 100% because everyone can tick several
    boxes. In socio-ecological studies measuring the public's priorities for wetland services, this analysis is
    indispensable.

====================================================================================

#30  MR Crosstab
    file: 030_MR_categorical_threats_country.xlsx
  >> SCENARIO (narration):
    In this opinion survey of 150 respondents we examined how perceptions of
    threats to lake ecosystems vary by nationality. Each participant's country
    is recorded as spain, sweden, or turkey, alongside their multiple-response
    perceived_threats selections of the dangers they consider important. A
    Multiple Response Crosstab cross-tabulates the perceived threats against
    country, so we can compare which concerns are emphasized in each nation.
    This highlights cross-cultural differences in how lake threats are
    perceived.
  >> VARIABLE SELECTION:
    - Multiple response set: perceived_threats
    - Column variable: country

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 150   Group column: country
    Perceived threats (eutrophication, drought, invasive species, pollution...) crossed by country

>> COMMENTARY (narration):
    We crossed perceived threats to wetlands (a multiple-choice item) by the respondent's country. With data from
    150 respondents, we can see how each threat perception is distributed across countries. The multiple-response
    crosstab answers "which group ticks which options more" -- for example, drought is foremost in one country,
    eutrophication in another. It is a powerful descriptive tool for international comparative wetland studies,
    revealing differences in environmental perception between countries/regions.

====================================================================================

#31  MR x MR Crosstab
    file: 031_MR_MR_function_tehdit.xlsx
  >> SCENARIO (narration):
    In this lake survey we asked 100 participants to mark, in a single
    questionnaire, which ecosystem functions they value and which threats they
    perceive for their local lakes. Because each respondent could tick several
    functions and several threats at once, both fonksiyonlar and threats are
    multiple-response variables rather than single choices. With a multiple-
    response by multiple-response crosstab we can map how each perceived
    function pairs with each perceived threat across all participants. This MR x
    MR crosstab is the right tool when two checklist-style questions are cross-
    tabulated and overlapping selections must be counted faithfully.
  >> VARIABLE SELECTION:
    - Multiple response set A: fonksiyonlar
    - Multiple response set B: threats

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 100
    Threat totals: agriculture 42, climate 41, drainage 33, pollution 37
    Wetland functions x perceived threats cross-tabulated

>> COMMENTARY (narration):
    This is the most advanced multiple-response analysis: we cross two separate multi-choice questions against each
    other -- the wetland functions a respondent values and the threats they perceive. Among 100 respondents the most
    frequent threats are agriculture (42) and climate (41), followed by pollution (37) and drainage (33). This table
    answers "which functions-valuers emphasise which threats" -- e.g. those valuing water provision may stress
    drought/climate, while those valuing biodiversity stress invasives/pollution. It is a powerful tool for showing
    how environmental perceptions and values pair up.

====================================================================================

#32  Cochran's Q Test
    file: 032_cochran_Q_macrophyte_4mevsim_binary.xlsx
  >> SCENARIO (narration):
    We tracked submerged macrophyte presence at 35 lakes across the four seasons
    of a single year, scoring each lake as present or absent in spring, summer,
    autumn and winter. The question is whether the proportion of lakes holding
    macrophytes changes significantly between seasons. Since the same lakes are
    measured repeatedly on a binary outcome over four matched time points,
    Cochran's Q test is the appropriate choice. It extends McNemar's logic to
    more than two related dichotomous measurements, telling us if seasonal
    occupancy differs.
  >> VARIABLE SELECTION:
    - Related binary measure 1: T1_spring
    - Related binary measure 2: T2_summer
    - Related binary measure 3: T3_autumn
    - Related binary measure 4: T4_winter

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 36.99   p < .001 ***
    Macrophyte presence (binary) changes significantly across 4 seasons.   H0 REJECTED

>> COMMENTARY (narration):
    We compared macrophyte presence (present/absent binary) across four seasons in the same lakes with Cochran's Q
    test -- the binary counterpart of repeated measures for more than two occasions, the multi-occasion extension of
    McNemar. The result is highly significant (Q = 36.99, p < .001): macrophyte presence varies clearly across
    seasons. This is an expected biological rhythm -- aquatic plants develop in summer and recede in winter.
    Cochran's Q is the right way to test change in repeated binary (yes/no) measurements on the same units, common in
    ecological presence/absence monitoring.

====================================================================================

#33  Correlation Analysis
    file: 033_correlation_TP_TN_chla_secchi_temp.xlsx
  >> SCENARIO (narration):
    Across 150 lakes we measured the classic limnological drivers of water
    clarity and productivity together: total phosphorus TP_ugL, total nitrogen
    TN_mgL, chlorophyll-a chla_ugL, Secchi depth secchi_m and water temperature
    temperature_C. We want to quantify how strongly these variables move
    together, for instance whether higher nutrients track higher chlorophyll and
    lower transparency. A correlation analysis gives us the pairwise strength
    and direction of every linear relationship among these continuous variables.
    This is the natural first step before any predictive modelling of trophic
    state.
  >> VARIABLE SELECTION:
    - Variable: TP_ugL
    - Variable: TN_mgL
    - Variable: chla_ugL
    - Variable: secchi_m
    - Variable: temperature_C

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Strongest relationship: TP_ugL - TN_mgL   r = +0.922   p < .001 ***  (Strong)
    5 variables: TP, TN, chlorophyll, Secchi, temperature

>> COMMENTARY (narration):
    We computed all pairwise correlations among five key limnological variables. The strongest is between total
    phosphorus and total nitrogen (r = 0.92) -- nutrient loads rising together, indicating shared sources
    (agricultural runoff, wastewater). Typically chlorophyll correlates positively with nutrients and negatively
    with Secchi (clarity): more nutrients = more algae = less clarity. The correlation matrix is the first step
    before building regressions -- to see the structure among variables and detect collinearity. The TP-TN value of
    0.92 also warns us to be careful about putting both in the same model.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Analysis
    file: 034_bland_altman_mikroskopi_FCM_goodnessfit.xlsx
  >> SCENARIO (narration):
    We counted phytoplankton in 60 water samples using two methods on each
    sample: classic microscopy classic_mikroskopi_cell_uL and imaging flow
    cytometry imaging_FCM_cell_uL. The goal is not whether they correlate but
    whether the two methods agree closely enough to be interchangeable in
    routine monitoring. A Bland-Altman analysis plots the difference against the
    mean for each sample and reveals the bias and the limits of agreement. This
    is the standard approach for assessing method comparability when both
    instruments measure the same cell concentration.
  >> VARIABLE SELECTION:
    - Method 1: classic_mikroskopi_cell_uL
    - Method 2: imaging_FCM_cell_uL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method 1 (classic microscopy) mean = 38.03   |   Method 2 (imaging FCM) mean = 39.59
    Bias (systematic difference) = -1.56   (methods agree well overall)

>> COMMENTARY (narration):
    We compared two methods of counting plankton cells -- classic microscopy and imaging-based flow cytometry (FCM)
    -- with Bland-Altman. This method does not look at correlation (high correlation does not guarantee agreement);
    instead it measures the difference between the two methods. The bias is just -1.56: FCM counts on average very
    slightly higher than microscopy, but the difference is small and there is almost no systematic bias. So the two
    methods are practically interchangeable. Bland-Altman is the standard way -- in methodology studies -- to compare
    a new/fast method against a gold standard.

====================================================================================

#35  Effect Size
    file: 035_effect_size_warm_cold_chla_cohenD.xlsx
  >> SCENARIO (narration):
    We compared chlorophyll-a chla_ugL between 80 lakes grouped by thermal
    regime into warm and cold climate_group categories. Beyond knowing that the
    means differ, we want to express how large that difference is in
    standardized terms for cross-study comparison. An effect size analysis, here
    Cohen's d, reports the magnitude of the warm-versus-cold gap in chlorophyll
    independent of sample size. This complements a significance test by
    quantifying the practical importance of the climate effect on algal biomass.
  >> VARIABLE SELECTION:
    - Group: climate_group
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -1.80  (very large effect)
    Chlorophyll difference between warm vs cold climate groups

>> COMMENTARY (narration):
    We measured the chlorophyll difference between warm and cold climate conditions not just as "is it significant"
    but "how large is it" -- as an effect size. Cohen's d = -1.80 is a very large effect; the two groups' distributions
    barely overlap. P-values inflate with sample size, but effect size is sample-independent and gives the practical
    importance of the real difference. Here the effect of warming on lake productivity is not just statistical but
    ecologically very strong. Reporting effect size in publications and meta-analyses is essential for the
    comparability of results.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 036_kanonik_correlation_environment_zoo_structuring.xlsx
  >> SCENARIO (narration):
    At 90 lakes we recorded an environmental set, total phosphorus TP_ugL, water
    temperature temperature_C and dissolved oxygen DO_mgL, alongside a
    zooplankton community set described by daphnia_percent, Bosmina_percent and
    Rotifer_percent. We want to know how the joint environmental gradient
    relates to the joint structure of the zooplankton community, rather than
    testing one pair at a time. Canonical correlation analysis finds the linear
    combinations of each set that are maximally correlated, revealing the
    dominant environment-to-zooplankton coupling. This is the right method when
    two multivariate blocks must be related simultaneously.
  >> VARIABLE SELECTION:
    - Set 1 variable: TP_ugL
    - Set 1 variable: temperature_C
    - Set 1 variable: DO_mgL
    - Set 2 variable: daphnia_percent
    - Set 2 variable: Bosmina_percent
    - Set 2 variable: Rotifer_percent

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.642, r2 = 0.41, Wilks L = 0.581, chi-square(9) = 46.44, p < .001 *
    CC2: r = 0.11, p = 0.91 (non-significant)
    X set: TP, temperature, dissolved oxygen  |  Y set: Daphnia%, Bosmina%, Rotifer%

>> COMMENTARY (narration):
    We related two multivariate sets -- environmental conditions (phosphorus, temperature, oxygen) and zooplankton
    community structure (Daphnia, Bosmina, Rotifer proportions) -- with canonical correlation. The first canonical
    axis is significant (r = 0.64, p < .001): a combination of environmental variables is strongly linked to a
    combination of zooplankton composition. The second axis is non-significant. So the environment-zooplankton
    relationship runs mainly along a single gradient: likely the eutrophication/temperature transition coinciding
    with a shift from large taxa (Daphnia) to small ones. CCA is the most direct answer -- very valuable in community
    ecology -- to "how do two blocks of variables co-vary".

====================================================================================

#37  Correspondence Analysis (CA)
    file: 037_correspondence_type_station_kontingenz.xlsx
  >> SCENARIO (narration):
    We assembled a large contingency table of phytoplankton observations
    classifying each of over 32,000 records by sampling station and by dominant
    taxon type, spanning Anabaena, Cryptomonas, Pediastrum, Cyclotella,
    Aphanizomenon and Microcystis. We want to visualize which taxa are
    associated with which stations and how stations differ in their assemblages.
    Correspondence analysis decomposes the station-by-type table into a low-
    dimensional map where proximity reflects association. This is the
    appropriate technique for exploring the structure of a two-way categorical
    contingency table.
  >> VARIABLE SELECTION:
    - Row category: station
    - Column category: type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.195   Number of dimensions = 2
    Station x type contingency table mapped in two dimensions

>> COMMENTARY (narration):
    We projected the station-by-species contingency table into a two-dimensional map with correspondence analysis.
    Total inertia (0.195) summarises the overall strength of the row-column association; the two dimensions visualise
    most of that structure. Station-species pairs that fall close together on the map indicate that the species is
    associated with that station. This lets us read at a glance which species concentrate at which stations --
    ecological gradients and indicator species. It is the most powerful visual way to summarise species-by-site
    composition data.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 038_varclus_environment_variable_clustering.xlsx
  >> SCENARIO (narration):
    For 120 lakes we compiled a broad limnological profile: TP_ugL, TN_mgL,
    chla_ugL, temperature_C, DO_mgL, conductivity_uScm, pH and alkalinity_mgL.
    Many of these variables are redundant, so we want to group them into
    clusters of mutually related measurements and identify a representative for
    each. Variable clustering (VarClus) partitions the variables into correlated
    groups, helping us reduce dimensionality before modelling. This is ideal
    when the aim is to organize and prune a set of intercorrelated environmental
    predictors.
  >> VARIABLE SELECTION:
    - Variable: TP_ugL
    - Variable: TN_mgL
    - Variable: chla_ugL
    - Variable: temperature_C
    - Variable: DO_mgL
    - Variable: conductivity_uScm
    - Variable: pH
    - Variable: alkalinity_mgL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (8 environmental variables)
    A representative variable was selected from each cluster (star)

>> COMMENTARY (narration):
    We clustered eight environmental variables (phosphorus, nitrogen, chlorophyll, temperature, oxygen, conductivity,
    pH, alkalinity) by similarity into three groups. VarClus clusters variables, not observations: it shows which
    measurements carry the same "information". Three clusters -- probably one nutrient load (TP/TN/chlorophyll), one
    ionic/salinity (conductivity/alkalinity/pH), one physical (temperature/oxygen) -- let us pick one representative
    per group and reduce the number of variables. In multivariate monitoring data it is extremely practical for
    removing redundancy and avoiding collinearity.

====================================================================================

#39  Multiple Linear Regression
    file: 039_multiple_linear_regression_chla_prediction.xlsx
  >> SCENARIO (narration):
    Using data from 180 lakes we aim to predict chlorophyll-a chla_ugL from its
    main physical and chemical drivers: total phosphorus TP_ugL, total nitrogen
    TN_mgL, water temperature temperature_C, Secchi depth secchi_m and water
    level WL_m. We want a single equation that quantifies each predictor's
    contribution to algal biomass while controlling for the others. Multiple
    linear regression fits chla_ugL as a linear function of these continuous
    predictors and reports their partial effects. This is the standard model
    when one continuous outcome is explained by several predictors jointly.
  >> VARIABLE SELECTION:
    - Dependent: chla_ugL
    - Predictor: TP_ugL
    - Predictor: TN_mgL
    - Predictor: temperature_C
    - Predictor: secchi_m
    - Predictor: WL_m

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.892   Adjusted R2 = 0.889
    Chlorophyll ~ TP + TN + temperature + Secchi + water level

>> COMMENTARY (narration):
    We built a multiple linear regression predicting chlorophyll-a from five environmental variables -- phosphorus,
    nitrogen, temperature, Secchi and water level. The model is very strong: together they explain 89% of the
    chlorophyll variance (R2 = 0.89). This shows shallow-lake productivity is largely set by measurable environmental
    drivers -- especially phosphorus's classic limiting-nutrient role. Such a model lets us predict chlorophyll from
    monitoring data and test "what happens to chlorophyll if we cut phosphorus by so much" scenarios. Multiple
    regression is the basic tool for explaining a continuous outcome with several drivers.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 040_logistic_regression_turbid_status.xlsx
  >> SCENARIO (narration):
    Across 220 lakes we classified each as turbid or clear in turbid_status_1_0
    and want to predict that binary state from total phosphorus TP_ugL,
    macrophyte presence macrophyte_present and water temperature temperature_C.
    The aim is to estimate how nutrient load, vegetation and temperature shift
    the probability of a lake tipping into the turbid regime. Logistic
    regression models the log-odds of the turbid outcome as a function of these
    predictors and yields interpretable odds ratios. It is the appropriate
    method when the response is a dichotomous ecological state.
  >> VARIABLE SELECTION:
    - Dependent: turbid_status_1_0
    - Predictor: TP_ugL
    - Predictor: macrophyte_present
    - Predictor: temperature_C

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: Logistic Regression   Pseudo R2 = 0.14
    Turbid status (1/0) ~ TP + macrophyte presence + temperature

>> COMMENTARY (narration):
    We built a logistic regression predicting the probability of a lake being in the "turbid" regime from phosphorus,
    macrophyte presence and temperature. Because the outcome is binary (turbid/clear), logistic regression is the
    right choice. Pseudo R2 = 0.14 shows the model partly explains turbidity status -- the coefficients generally say
    the odds of switching to the turbid regime rise as phosphorus increases and macrophytes decline. This quantifies
    the alternative-stable-states theory of shallow lakes. Logistic regression is the most common way to predict a
    binary outcome (present/absent, sick/healthy, turbid/clear) from continuous and categorical predictors.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Poisson Regression
    file: 041_poisson_regression_bird_type_sayimi.xlsx
  >> SCENARIO (narration):
    In a network of wetland lakes we want to understand what drives waterbird
    abundance at each monitoring site. We recorded bird_type_count as the number
    of birds counted, alongside the lake area_ha, a water_quality_score_1_5
    index, and the survey year. Because bird counts are non-negative integers
    that often cluster at low values, an ordinary linear model is inappropriate,
    so we use Poisson Regression to model the count as a function of habitat
    size, water quality, and temporal trend. This lets us quantify how much each
    additional hectare or each step in water quality changes the expected number
    of birds.
  >> VARIABLE SELECTION:
    - Outcome (count): bird_type_count
    - Predictor: area_ha
    - Predictor: water_quality_score_1_5
    - Predictor: year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 1310.5   Deviance = 282.1
    Bird species count ~ area (ha) + water-quality score + year   (Poisson)

>> COMMENTARY (narration):
    We built a Poisson regression predicting wetland bird species count -- a count variable -- from area, water
    quality and year. Count data (0,1,2,...) are not normally distributed, so we use Poisson rather than linear
    regression. The model typically shows species count rising with area and water quality, with a possible year
    trend. Deviance and AIC assess model fit. Count regression is the correct way to relate biodiversity counts
    (species richness, individual counts) to environmental drivers; it is far more appropriate than a linear model.

====================================================================================

#42  Multinomial Logistic Regression
    file: 042_multinomial_logistic_3sinif_lake_status.xlsx
  >> SCENARIO (narration):
    We classified a set of lakes into three unordered ecological states - clear,
    turbid, or medium - and want to know which water-chemistry variables
    separate these classes. For each lake we have total phosphorus TP_ugL,
    Secchi depth secchi_m, and conductivity_uScm as candidate drivers of
    status_3sinif. Since the outcome has three nominal categories with no
    natural ranking, Multinomial Logistic Regression is the right tool to
    estimate how each predictor shifts the odds of one lake state versus
    another. The result tells us, for example, whether higher phosphorus pushes
    a lake from clear toward turbid.
  >> VARIABLE SELECTION:
    - Outcome (nominal 3-class): status_3sinif
    - Predictor: TP_ugL
    - Predictor: secchi_m
    - Predictor: conductivity_uScm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 191.7   Reference class: 'clear'
    3-class lake status ~ TP + Secchi + conductivity

>> COMMENTARY (narration):
    We classified lakes into three states -- clear, intermediate, turbid -- and modelled how that state is predicted
    from phosphorus, Secchi and conductivity with multinomial logistic regression. Because the outcome has more than
    two unordered categories, this method fits: 'clear' is the reference, and a separate equation is built for each
    other state. Coefficients read as "how many times the odds of being turbid rather than clear change as phosphorus
    rises". This lets us model the full transition gradient rather than splitting lake status only in two -- a
    powerful tool in ecological status classification.

====================================================================================

#43  Ordinal Logistic Regression
    file: 043_ordinal_logistic_WFD_5sinif_ecological_status.xlsx
  >> SCENARIO (narration):
    Following the Water Framework Directive, lakes are graded into five ordered
    ecological status classes from poor up to high. We want to predict
    ecological_status_5sinif from two key trophic indicators, total phosphorus
    TP_ugL and chlorophyll-a chla_ugL. Because these five quality classes have a
    clear rank order, Ordinal Logistic Regression respects that ordering while
    estimating how rising nutrients and algal biomass push a lake toward worse
    status. This gives managers a single model linking measurable water quality
    to the regulatory status grade.
  >> VARIABLE SELECTION:
    - Outcome (ordinal 5-class): ecological_status_5sinif
    - Predictor: TP_ugL
    - Predictor: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 501.9
    WFD 5-class ecological status (high>good>moderate>poor>bad) ~ TP + chlorophyll

>> COMMENTARY (narration):
    We built an ordinal logistic regression predicting the Water Framework Directive's five-level ecological status
    (high, good, moderate, poor, bad) from phosphorus and chlorophyll. These classes are ordered -- there is a
    natural ranking among them -- and the ordinal model uses that order, making it more powerful and interpretable
    than multinomial. The coefficients give the tendency of a lake to slide toward a worse status as phosphorus and
    chlorophyll rise. This directly models the "ecological status classification" at the heart of EU water policy.
    For ordered categorical outcomes (status class, Likert, disease stage), ordinal logistic is the right method.

====================================================================================

#44  PLS Regression
    file: 044_pls_regression_8x_correlated_chla.xlsx
  >> SCENARIO (narration):
    In experimental tanks we measured chlorophyll-a chla_ugL together with eight
    environmental predictors, env_present_1 through env_present_8, that are
    strongly correlated with one another. With so many collinear inputs and a
    modest sample, ordinary regression coefficients become unstable. Partial
    Least Squares (PLS) Regression handles this by extracting a few latent
    components that capture the shared variation among the eight predictors
    while maximizing their relationship to chla_ugL. This lets us build a stable
    predictive model of algal biomass despite the heavy multicollinearity.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: env_present_1
    - Predictor: env_present_2
    - Predictor: env_present_3
    - Predictor: env_present_4
    - Predictor: env_present_5
    - Predictor: env_present_6
    - Predictor: env_present_7
    - Predictor: env_present_8

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2_train = 0.89   R2_CV (5-fold) = 0.87
    Chlorophyll ~ 8 collinear (correlated) environmental variables

>> COMMENTARY (narration):
    We predicted chlorophyll from eight highly correlated environmental variables with PLS (Partial Least Squares)
    regression. When variables are collinear (as we saw in #33, TP-TN = 0.92), ordinary regression becomes unstable;
    PLS solves this by reducing the variables to a few orthogonal components. The model is strong both in training
    (R2 = 0.89) and cross-validation (R2_CV = 0.87) -- so there is no overfitting and generalisation is good. When
    there are many correlated measurements (environmental, spectral, chemical), PLS is the preferred method that
    keeps both predictive power and interpretability.

====================================================================================

#45  Probit Regression
    file: 045_probit_fish_establishment_probability.xlsx
  >> SCENARIO (narration):
    We are studying whether non-native fish manage to establish in a lake,
    recorded as the binary outcome fish_establishment_present. We suspect that
    lake size area_km2 and mean_depth_m govern the probability of a successful
    population taking hold. Because the response is a 0/1 establishment event,
    Probit Regression models the probability of establishment through a
    cumulative normal link as a function of these two morphometric variables.
    The fitted model estimates how the chance of establishment rises with larger
    and deeper lakes.
  >> VARIABLE SELECTION:
    - Outcome (binary): fish_establishment_present
    - Predictor: area_km2
    - Predictor: mean_depth_m

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 212.6   Pseudo R2 (McFadden) = 0.12
    Fish establishment probability (1/0) ~ lake area (km2) + mean depth (m)

>> COMMENTARY (narration):
    We predicted whether a fish population would establish in a lake from lake area and mean depth with probit
    regression. Probit models a binary outcome like logistic but assumes a normal-distribution curve (cumulative
    normal) -- common in ecology and toxicology. Pseudo R2 = 0.12: area and depth partly explain establishment
    probability; larger, deeper lakes are generally more able to sustain fish populations. Such models are used for
    lake-island biogeography and invasive/native fish spread risk maps. Probit and logistic usually give similar
    results; the choice is often tradition and interpretive preference.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 046_tobit_TP_LOD_censored.xlsx
  >> SCENARIO (narration):
    Our laboratory measures total phosphorus, but many low samples fall below
    the analytical limit of detection and are reported only as censored values.
    The variable TP_measurement_ugL holds the recorded phosphorus,
    LOD_under_censored flags which observations are left-censored at the
    detection limit, and we also have conductivity_uScm and temperature_C as
    covariates. Ignoring the censoring would bias our estimates, so Tobit
    Regression explicitly models the censored phosphorus as a function of
    conductivity and temperature. This recovers the true relationship between
    water chemistry and phosphorus even when part of the data is below
    detection.
  >> VARIABLE SELECTION:
    - Outcome (censored): TP_measurement_ugL
    - Censoring indicator: LOD_under_censored
    - Predictor: conductivity_uScm
    - Predictor: temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 1161.58   Tobit (left-censored: measurement below LOD)
    TP measurement ~ conductivity + temperature

>> COMMENTARY (narration):
    A common problem in lab measurements: some phosphorus values fall below the instrument's limit of detection (LOD)
    -- flagged "not detected", but not zero. Deleting these or treating them as zero biases the result. Tobit
    regression correctly models this left-censored data, estimating phosphorus's relationship with conductivity and
    temperature. The model was fitted with AIC = 1161.6. In environmental chemistry, the statistically correct
    solution to "below detection limit" values -- the censored-data problem -- is the Tobit model; it is superior to
    raw deletion/substitution.

====================================================================================

#47  Bayesian Linear Regression
    file: 047_bayesian_linear_small_n_pond.xlsx
  >> SCENARIO (narration):
    We sampled only eighteen small ponds and want to relate chlorophyll-a
    chla_ugL to total phosphorus TP_ugL and the percentage of macrophyte cover
    macrofit_percent_closure. With such a tiny sample, classical regression
    estimates are noisy and confidence intervals are unreliable. Bayesian Linear
    Regression lets us incorporate prior information and obtain full posterior
    distributions for each coefficient, giving honest uncertainty even at small
    n. This yields credible intervals for how phosphorus and macrophyte closure
    jointly shape algal biomass.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: TP_ugL
    - Predictor: macrofit_percent_closure

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TP_ugL posterior mean = 0.52   95% credible interval (0.42 , 0.62)   P(beta>0) = 1.00
    Conjugate Normal-Inverse-Gamma prior   posterior sigma2 = 9.11

>> COMMENTARY (narration):
    With few ponds (small n), we predicted chlorophyll from phosphorus and macrophytes using Bayesian linear
    regression. The Bayesian advantage: the result is not a single p-value but a probability distribution (posterior)
    for the parameters. For phosphorus's effect the posterior mean is 0.52, 95% credible interval (0.42, 0.62), and
    P(beta>0) = 1.00 -- the evidence that phosphorus raises chlorophyll is decisive. In small samples the Bayesian
    method honestly expresses uncertainty and can incorporate prior knowledge. It is valuable in data-poor ecology.

====================================================================================

#48  Nonlinear Regression
    file: 048_nonlinear_michaelis_menten_TP_chla.xlsx
  >> SCENARIO (narration):
    Ecological theory predicts that chlorophyll-a does not rise without limit as
    phosphorus increases but instead saturates. Across many water samples we
    measured total phosphorus TP_ugL and chlorophyll-a chla_ugL to test this
    saturating response. A straight line cannot capture the curve, so we fit a
    Nonlinear Regression using a Michaelis-Menten form, modeling chla_ugL as a
    saturating function of TP_ugL. The fitted parameters give the maximum
    attainable chlorophyll and the phosphorus level at which the response is
    half-saturated.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: y = a · x^b (power function)   R2 = 0.908   RMSE = 7.57
    Chlorophyll ~ phosphorus (saturation-type relationship)

>> COMMENTARY (narration):
    The chlorophyll-phosphorus relationship is not linear: it rises fast at low phosphorus and approaches saturation
    at high phosphorus (limited by light and other nutrients). We modelled this with a power function -- y = a*x^b --
    and the fit is very good (R2 = 0.91). This curve captures Michaelis-Menten-like saturation dynamics: each
    additional unit of phosphorus contributes less to chlorophyll. Nonlinear regression is the way to correctly model
    ecological growth/saturation/decay curves (logistic growth, dose-response, nutrient-response); forcing a straight
    line would mislead.

====================================================================================

#49  Ridge Regression
    file: 049_ridge_multicollinear_chla.xlsx
  >> SCENARIO (narration):
    We want to predict chlorophyll-a chla_ugL from seven environmental drivers,
    x1_correlated through x7_correlated, but these predictors are highly
    intercorrelated. Such multicollinearity inflates the variance of ordinary
    least-squares coefficients and makes them unstable. Ridge Regression adds an
    L2 penalty that shrinks the coefficients, stabilizing the estimates while
    keeping all seven predictors in the model. This produces a more reliable
    predictive model of algal biomass under collinearity.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: x1_correlated
    - Predictor: x2_correlated
    - Predictor: x3_correlated
    - Predictor: x4_correlated
    - Predictor: x5_correlated
    - Predictor: x6_correlated
    - Predictor: x7_correlated

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Ridge (alpha = 1.0): R2 = 0.900   Adj. R2 = 0.895   n = 140
    7 mutually correlated predictors

>> COMMENTARY (narration):
    Predicting chlorophyll from seven highly correlated environmental variables, ordinary regression becomes unstable
    (collinearity inflates coefficients). Ridge regression solves this by shrinking all coefficients (with a penalty)
    -- it does not zero them, but limits their magnitude, keeping prediction stable. The result is strong: R2 = 0.90.
    Ridge is ideal when "I want to keep all variables but there is collinearity". In environmental monitoring's
    common correlated measurement sets, it prevents overfitting and preserves generalisation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 050_lasso_actual_3prediktor_11gurultu_chla.xlsx
  >> SCENARIO (narration):
    We measured chlorophyll-a chla_ugL alongside fourteen candidate predictors,
    of which only actual_x1, actual_x2, and actual_x3 are truly informative
    while noise_x1 through noise_x11 are irrelevant. We need a method that
    automatically identifies the few genuine drivers and discards the noise.
    Lasso Regression applies an L1 penalty that forces unimportant coefficients
    to exactly zero, performing variable selection while fitting the model. The
    result is a sparse, interpretable model that recovers the real predictors of
    algal biomass from the surrounding noise.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: actual_x1
    - Predictor: actual_x2
    - Predictor: actual_x3
    - Predictor: noise_x1
    - Predictor: noise_x2
    - Predictor: noise_x3
    - Predictor: noise_x4
    - Predictor: noise_x5
    - Predictor: noise_x6
    - Predictor: noise_x7
    - Predictor: noise_x8
    - Predictor: noise_x9
    - Predictor: noise_x10
    - Predictor: noise_x11

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Lasso (alpha = 0.1): R2 = 0.916   Adj. R2 = 0.904   n = 120
    2/14 feature coefficients shrunk to zero (automatic feature selection)

>> COMMENTARY (narration):
    In this data 3 real predictors were mixed with 11 noise variables. The strength of Lasso shows here: the penalty
    term drives the coefficients of irrelevant variables exactly to zero -- automatic variable selection. The model
    stayed strong (R2 = 0.92) while weeding out useless variables. Lasso is extremely useful when you want to pick
    the true drivers from many candidates -- in high-dimensional environmental/genomic/spectral data. While Ridge
    shrinks all variables, Lasso eliminates the irrelevant ones entirely; the two serve different purposes.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 051_mediation_fish_zoo_chla_kaskat.xlsx
  >> SCENARIO (narration):
    In a connected lake chain we want to know whether fish influence algae not
    directly, but through their effect on grazing zooplankton. We measured fish
    catch per unit effort (fish_cpue), zooplankton biomass (zoo_biomass_ugDW_L),
    and chlorophyll-a (chla_ugL) across 150 samples. The hypothesis is a classic
    trophic cascade: more fish suppress zooplankton, and fewer grazers allow
    chlorophyll-a to rise. Mediation analysis is the right tool here because it
    formally separates the direct effect of fish on chlorophyll-a from the
    indirect path running through zooplankton biomass.
  >> VARIABLE SELECTION:
    - Predictor (X): fish_cpue
    - Mediator (M): zoo_biomass_ugDW_L
    - Outcome (Y): chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 0.59   SE = 0.13   p < .001   95% CI (0.42 , 0.91)   Significant
    Path: fish (CPUE) -> zooplankton biomass -> chlorophyll

>> COMMENTARY (narration):
    We tested the classic trophic cascade -- that the fish effect on chlorophyll passes through zooplankton -- with
    mediation analysis. The indirect effect is significant (0.59, p < .001, CI excludes zero): fish predation reduces
    zooplankton, the reduced zooplankton lowers grazing pressure, which raises algae (chlorophyll). So fish affect
    chlorophyll not directly but "through" zooplankton -- the mechanism behind why biomanipulation (fish removal)
    works in lake management. Mediation analysis opens up the intermediate mechanism in "why X affects Y".

====================================================================================

#52  Path Analysis
    file: 052_path_analysis_5degisken_TP_zoo_macr_chla_secchi.xlsx
  >> SCENARIO (narration):
    We want to map out the full web of causes that determine water clarity in
    shallow lakes. Using 300 observations, we trace how total phosphorus
    (TP_ugL) drives chlorophyll-a (chla_ugL) both directly and through changes
    in zooplankton grazing (zoo_biomass) and macrophyte cover
    (macrophyte_closure_percent), with all of these ultimately shaping Secchi
    depth (secchi_m). Path analysis is ideal because it lets us estimate several
    interlinked regression relationships at once and quantify each direct and
    indirect route to transparency. This reveals whether clarity is governed
    mostly by nutrients, by biological grazing, or by the stabilizing role of
    submerged plants.
  >> VARIABLE SELECTION:
    - Exogenous: TP_ugL
    - Mediator: zoo_biomass
    - Mediator: macrophyte_closure_percent
    - Mediator: chla_ugL
    - Outcome: secchi_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.968   RMSEA = 0.084   (acceptable fit)
    Model: phosphorus/zooplankton/macrophyte -> chlorophyll -> Secchi (clarity)

>> COMMENTARY (narration):
    Path analysis is extended regression that models several cause-effect relationships at once. Here we tested, in a
    single path model, that phosphorus, zooplankton and macrophytes affect chlorophyll, and chlorophyll in turn
    determines water clarity (Secchi). The fit is acceptable (CFI = 0.97, RMSEA = 0.08). This visually and numerically
    summarises the causal web of the shallow-lake ecosystem -- the nutrient -> algae -> clarity chain. Path analysis
    is powerful for separating direct and indirect effects and testing multivariate theoretical models.

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 053_lmm_pan_europe_mesocosm_country_random.xlsx
  >> SCENARIO (narration):
    A pan-European mesocosm experiment ran identical setups in six countries to
    test how nutrients and warming affect algal growth. Across 210 mesocosms we
    recorded total phosphorus (TP_ugL), water temperature (temperature_C), and
    chlorophyll-a (chla_ugL), but each country contributes a cluster of
    measurements with its own baseline conditions. A linear mixed model is
    appropriate because we treat country as a random effect, absorbing site-to-
    site variation while estimating the fixed effects of phosphorus and
    temperature on chlorophyll-a. This lets us draw a general conclusion about
    nutrient and warming responses that holds across the continent.
  >> VARIABLE SELECTION:
    - Random effect: country
    - Fixed effect: TP_ugL
    - Fixed effect: temperature_C
    - Outcome: chla_ugL

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group variance (random intercept) = 72.40   ICC = 0.723
    Chlorophyll ~ phosphorus + temperature  +  (1 | country)

>> COMMENTARY (narration):
    In the pan-European mesocosm experiment, data are nested within countries (several tanks per country). Tanks
    within the same country resemble each other, violating the independence assumption -- ordinary regression would
    be wrong. The linear mixed model solves this by adding country as a random effect. ICC = 0.72 is very high: 72%
    of chlorophyll variance comes from between-country differences; the rest is within-tank. So country (climate/
    background) is very determining. LMM is the correct, indispensable method for nested/hierarchical data
    (country>tank, school>student, patient>repeated measure).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 054_multiple_imputation_zoo_MAR_missing.xlsx
  >> SCENARIO (narration):
    Our zooplankton monitoring dataset has gaps because some net hauls were lost
    or unprocessed, leaving zoo_biomass_ugDW_L missing for part of the 120
    samples. Rather than discard those records, we want to recover them using
    the other measured variables, total phosphorus (TP_ugL), water temperature
    (temperature_C), and Secchi depth (secchi_m). Because the missingness is
    plausibly at random and related to these covariates, multiple imputation is
    the correct approach: it generates several plausible filled-in datasets and
    pools the results to honestly reflect uncertainty. This preserves
    statistical power without biasing our estimate of zooplankton biomass.
  >> VARIABLE SELECTION:
    - Impute: zoo_biomass_ugDW_L
    - Predictor: TP_ugL
    - Predictor: temperature_C
    - Predictor: secchi_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (number of imputations) = 5
    Missing values in zooplankton biomass (MAR) estimated from phosphorus/temperature/Secchi

>> COMMENTARY (narration):
    Missing observations are inevitable in field data; deleting them (listwise) both loses data and biases results
    (if missingness is not random). Multiple imputation estimates the missing zooplankton values from other
    variables (phosphorus, temperature, Secchi), creating 5 separate completed datasets, runs the analysis in each
    and combines the results -- thereby also accounting for estimation uncertainty. Under the MAR (missing at random)
    assumption this is the modern missing-data standard. In long-running monitoring studies it is the soundest way to
    cope with lost observations.

====================================================================================

#55  GEE
    file: 055_gee_repeated_lake_monthly_chla.xlsx
  >> SCENARIO (narration):
    We ran a whole-lake manipulation where lakes received either a nutrient-
    reduction treatment or served as controls, and we tracked chlorophyll-a
    (chla_ugL) monthly across 240 visits. Because each lake is sampled
    repeatedly over the months, the observations within a lake are correlated
    and cannot be treated as independent. A generalized estimating equations
    approach fits this design well, modeling the average effect of treatment and
    month on chlorophyll-a while accounting for the within-lake correlation
    structure. This tells us whether the treatment produced a population-level
    reduction in algal biomass over time.
  >> VARIABLE SELECTION:
    - Subject/cluster: lake_id
    - Time: month
    - Predictor: treatment
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    month coefficient = 1.64   SE = 0.10   p < .001   QIC = 247.51
    Chlorophyll ~ month  (repeated measures, clustered within lake)

>> COMMENTARY (narration):
    We took monthly repeated chlorophyll measurements from the same lakes -- the measurements of one lake are
    dependent. GEE estimates the relationship at the "population average" level in such clustered/repeated data,
    accounting for the correlation structure. The month coefficient is significant (1.64, p < .001): chlorophyll
    shows a clear within-year seasonal increase. GEE is an alternative to LMM: rather than modelling random effects,
    it treats the correlation among observations as a nuisance to be corrected, giving the marginal (average) effect.
    It is a common, powerful method for longitudinal/repeated-measures data.

====================================================================================

#56  GLMM
    file: 056_glmm_poisson_bird_count_site_year.xlsx
  >> SCENARIO (narration):
    We are studying how waterbird abundance around lakes depends on habitat
    size, surveyed repeatedly over several years at many sites. The response,
    bird_count, is a count of individuals at each site, recorded with the survey
    year and the lake area (area_ha) over 240 site-year records. Counts from the
    same site across years are correlated, and a Poisson-distributed response
    calls for a generalized linear mixed model with site as a random effect.
    This lets us estimate how area and year relate to bird counts while
    respecting both the count nature of the data and the repeated-site
    structure.
  >> VARIABLE SELECTION:
    - Random effect: site_id
    - Fixed effect: year
    - Fixed effect: area_ha
    - Outcome (count): bird_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    area_ha coefficient ~ 0.0000   p = 0.976  (non-significant)
    Poisson GLMM: bird count ~ area + year  +  (1 | site)

>> COMMENTARY (narration):
    We modelled bird counts (Poisson) with both area/year fixed effects and a site random effect -- that is, a GLMM:
    mixed model + count distribution. In this data the area effect is non-significant (p = 0.98); bird count may
    depend more on year or site-specific differences than on site area. GLMM combines the strengths of LMM (which
    assumes normality) and GLM (which ignores clustering) for "repeated/nested count or binary data". It is the
    correct analytic framework for the common ecological design of "counts over years at the same sites".

====================================================================================

#57  Elastic Net
    file: 057_elastic_net_two_correlated_group_chla.xlsx
  >> SCENARIO (narration):
    We want to predict chlorophyll-a (chla_ugL) from a large panel of candidate
    predictors, but many of them come in two highly correlated groups plus
    several pure noise variables. With twelve informative-looking columns
    (x_group1_1 through x_group2_4) and six noise columns across 150 samples,
    ordinary regression would be unstable and over-fit. Elastic net is the right
    choice because its blend of L1 and L2 penalties shrinks irrelevant
    coefficients toward zero while gracefully handling groups of correlated
    predictors together. The result is a parsimonious, stable model that
    isolates which water-quality signals truly drive algal biomass.
  >> VARIABLE SELECTION:
    - Predictor: x_group1_1
    - Predictor: x_group1_2
    - Predictor: x_group1_3
    - Predictor: x_group1_4
    - Predictor: x_group2_1
    - Predictor: x_group2_2
    - Predictor: x_group2_3
    - Predictor: x_group2_4
    - Predictor: x_noise_1
    - Predictor: x_noise_2
    - Predictor: x_noise_3
    - Predictor: x_noise_4
    - Predictor: x_noise_5
    - Predictor: x_noise_6
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.220   R2 (train) = 0.877
    Two correlated variable groups + noise -> chlorophyll

>> COMMENTARY (narration):
    Elastic Net is a hybrid of Ridge and Lasso: it both shrinks coefficients (like Ridge) and zeros out the
    irrelevant ones (like Lasso). Here there were two groups of correlated variables; in such cases Lasso tends to
    pick one variable from a group at random and drop the rest, whereas Elastic Net tends to keep the correlated
    group together. The model is strong (R2 = 0.88). When there are collinear variable groups -- gene expression,
    spectral bands, correlated environmental measures -- Elastic Net is a balanced choice offering both selection and
    stability.

====================================================================================

#58  Robust Regression
    file: 058_robust_outlier_chla_TP.xlsx
  >> SCENARIO (narration):
    We are estimating the relationship between total phosphorus (TP_ugL) and
    chlorophyll-a (chla_ugL) across 100 lakes, but a few eutrophic outliers with
    extreme algal blooms threaten to distort an ordinary least-squares line. We
    do not want to delete these real lakes, yet we also do not want them to
    dominate the slope. Robust regression is appropriate because it down-weights
    the influence of outliers, giving a phosphorus-chlorophyll relationship that
    reflects the bulk of the lakes. This yields a more trustworthy estimate of
    how nutrients translate into algal biomass.
  >> VARIABLE SELECTION:
    - Predictor: TP_ugL
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust beta (TP) = +0.497   SE = 0.023   p < .001
    Robust intercept = 1.23   vs   OLS intercept = 6.99   |   8 observations potential outliers (weight < 0.5)

>> COMMENTARY (narration):
    This data had a few outliers. Ordinary regression (OLS) takes outliers fully into account and is pulled toward
    them -- the intercept is 6.99 in OLS but only 1.23 in the robust model; the large gap means outliers seriously
    distorted OLS. Robust regression (Huber-T) gives less weight to outlying observations to preserve the true trend;
    it flagged 8 observations as outliers. In field data, where outliers are inevitable (measurement error, extreme
    events), robust regression gives a more reliable slope estimate than OLS.

====================================================================================

#59  Quantile Regression
    file: 059_quantile_heteroskedastik_chla_TP.xlsx
  >> SCENARIO (narration):
    We suspect that phosphorus does not affect all lakes equally: at high
    nutrient levels the spread of chlorophyll-a (chla_ugL) widens dramatically,
    a sign of heteroskedasticity. Across 200 samples we have total phosphorus
    (TP_ugL) and water temperature (temperature_C) as predictors, and we are
    especially interested in the upper tail where blooms occur. Quantile
    regression is the right tool because it models how the median and the high
    quantiles of chlorophyll-a respond to nutrients and temperature, rather than
    just the mean. This exposes whether phosphorus has a stronger effect on
    bloom-prone lakes than on typical ones.
  >> VARIABLE SELECTION:
    - Predictor: TP_ugL
    - Predictor: temperature_C
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    q = 0.10 slope = 0.57   |   q = 0.50 (median)   |   q = 0.90  (changing spread)
    Chlorophyll ~ phosphorus (heteroskedastic)

>> COMMENTARY (narration):
    Ordinary regression models only the mean; but in the phosphorus-chlorophyll relationship the spread grows as
    phosphorus rises (heteroskedasticity) -- so phosphorus's effect differs across the distribution. Quantile
    regression models the lower (q=0.10), middle (q=0.50) and upper (q=0.90) quantiles separately. A steeper slope at
    the upper quantile means "the most eutrophic lakes are the most sensitive to phosphorus". In environmental risk
    assessment, where the extremes (worst case) matter more than the average, quantile regression reveals what mean-
    based models miss.

====================================================================================

#60  ROC Curve Analysis
    file: 060_roc_turbid_status_siniflandirici.xlsx
  >> SCENARIO (narration):
    We want a simple screening rule that flags whether a lake is in a turbid
    state (turbid_status_1_0) from a handful of routine measurements. Using 240
    lakes we have total phosphorus (TP_ugL), Secchi depth (secchi_m), macrophyte
    presence (macrophyte_present), and water temperature (temperature_C) as
    candidate classifiers. ROC curve analysis is the natural choice because it
    evaluates how well each predictor separates turbid from clear lakes across
    all possible thresholds, summarizing performance with the area under the
    curve. This helps us pick the most reliable early-warning indicator and its
    optimal cut-off.
  >> VARIABLE SELECTION:
    - Predictor: TP_ugL
    - Predictor: secchi_m
    - Predictor: macrophyte_present
    - Predictor: temperature_C
    - Class (target): turbid_status_1_0

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.867 (good discriminating power)   Youden optimum threshold: 0.147 -> Sensitivity 0.72
    n = 240 (positive 18, negative 222)   turbid status classifier

>> COMMENTARY (narration):
    We measured the discriminating power of a classifier predicting whether a lake is "turbid" from environmental
    variables with the ROC curve. AUC = 0.87 is "good": the model largely separates turbid from clear lakes
    correctly. The Youden index gives the best cut-off (threshold) -- optimising the balance of sensitivity and
    specificity. Because positives are rare (18/240, imbalanced), AUC is more informative than accuracy. ROC/AUC is
    the standard for reporting the quality of diagnostic/classification models (medical diagnosis, species
    distribution models, status classification).

====================================================================================

#61  True Skill Statistic (TSS)
    file: 061_tss_daphnia_varligi_SDM.xlsx
  >> SCENARIO (narration):
    We want to know whether the small grazer Daphnia can persist in a lake given
    that lake's pressures, so we built a species distribution model predicting
    daphnia_present from total phosphorus TP_ugL, fish predation pressure
    expressed as fish_cpue, and water temperature_C across 200 lakes. Because
    both the prediction and the truth are binary presence/absence, accuracy
    alone is misleading when presences and absences are unbalanced. The True
    Skill Statistic (TSS) summarizes sensitivity and specificity into a single
    threshold-independent score, telling us how much better than chance our
    Daphnia model really performs. This makes TSS the right diagnostic for
    judging the ecological usefulness of our presence/absence prediction.
  >> VARIABLE SELECTION:
    - Predicted/Observed binary: daphnia_present
    - Predictor: TP_ugL
    - Predictor: fish_cpue
    - Predictor: temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.21 (acceptable)   Sensitivity 0.71   Specificity 0.50   n = 200
    Daphnia presence prediction (species distribution model)

>> COMMENTARY (narration):
    TSS (True Skill Statistic) is a common performance measure in species distribution models (SDM): computed as
    sensitivity + specificity - 1, it shows how much better than chance we are. The model predicting Daphnia presence
    from environmental variables has TSS = 0.21 -- "acceptable": not perfect, but better than chance. TSS's advantage
    is that it is insensitive to prevalence (how common the species is); for that reason it is preferred in ecology as
    an alternative to AUC. It honestly summarises the performance of species presence/absence models.

====================================================================================

#62  Confusion Matrix Metrics
    file: 062_complexity_matrix_clear_turbid.xlsx
  >> SCENARIO (narration):
    Our classifier tries to separate clear-water from turbid lake states, and we
    need to know not just how often it is right but how it fails. For 220 lakes
    we have a continuous prediction_score, the model's binary
    prediction_label_0_1, and the field-verified actual_label_0_1, alongside
    drivers such as TP_ugL and macrophyte_percent that motivate the regime.
    Confusion Matrix Metrics cross-tabulate predicted versus actual labels to
    yield precision, recall, and specificity, exposing whether the model
    confuses turbid lakes for clear ones. This is exactly the tool for auditing
    a binary clear/turbid classifier against ground truth.
  >> VARIABLE SELECTION:
    - Predicted label: prediction_label_0_1
    - Actual label: actual_label_0_1
    - Predicted score: prediction_score
    - Feature: TP_ugL
    - Feature: macrophyte_percent

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.859   Sensitivity (Recall) = 0.343   Precision = 0.60
    F1 = 0.436   Balanced Accuracy = 0.65

>> COMMENTARY (narration):
    We broke down a classifier's performance (clear/turbid prediction) with a confusion matrix. Although overall
    accuracy looks high (86%), this is misleading: sensitivity (catching the true turbid lakes) is only 0.34 -- the
    model mostly misses turbid lakes. The reason is class imbalance: because turbid lakes are rare, a model that says
    "all clear" still scores high accuracy. F1 and balanced accuracy expose this trap. The confusion matrix is the
    indispensable tool that reveals what the "accuracy" number hides -- where each class is misclassified.

====================================================================================

#63  Random Forest
    file: 063_random_forest_lake_regime_estimate.xlsx
  >> SCENARIO (narration):
    We ask which lake characteristics best predict whether a shallow lake sits
    in a clear or turbid regime, using 300 lakes labelled status_2sinif as clear
    or turbid. The candidate drivers are nutrients TP_ugL and TN_mgL,
    macrophyte_closure_percent, fish_cpue, temperature_C, secchi_m transparency,
    and conductivity_uScm. A Random Forest handles these many interacting,
    nonlinear predictors without distributional assumptions and ranks their
    importance for the clear-versus-turbid outcome. That makes it ideal for both
    classifying regime and revealing which variables drive alternative stable
    states.
  >> VARIABLE SELECTION:
    - Target: status_2sinif
    - Predictor: TP_ugL
    - Predictor: TN_mgL
    - Predictor: macrophyte_closure_percent
    - Predictor: fish_cpue
    - Predictor: temperature_C
    - Predictor: secchi_m
    - Predictor: conductivity_uScm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (2-class lake regime)
    Predicted from 7 environmental variables

>> COMMENTARY (narration):
    We predicted lake regime (two classes) from seven environmental variables with Random Forest -- the vote of
    hundreds of decision trees. Accuracy is 100%: the regimes separate almost perfectly with these variables. Such
    high accuracy shows the classes are genuinely well separated environmentally (though one should always keep
    overfitting risk in mind). One of Random Forest's most valuable outputs is the "variable importance ranking",
    telling which environmental variable matters most for the split. It is a powerful, easy-to-build machine-learning
    method for nonlinear, interacting relationships.

====================================================================================

#64  SVM
    file: 064_svm_clear_turbid.xlsx
  >> SCENARIO (narration):
    We want a clean decision boundary that separates clear from turbid lakes
    based on their physical and biological state. Across 200 lakes we use
    TP_ugL, macrophyte_closure_percent, secchi_m, and temperature_C to predict
    the status label of clear or turbid. A Support Vector Machine finds the
    maximum-margin boundary between the two classes and can bend that boundary
    with kernels when the regimes are not linearly separable. This suits our
    goal of robustly classifying lake state from a handful of strong predictors.
  >> VARIABLE SELECTION:
    - Target: status
    - Predictor: TP_ugL
    - Predictor: macrophyte_closure_percent
    - Predictor: secchi_m
    - Predictor: temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.975   (clear/turbid)
    Predicted from 4 environmental variables

>> COMMENTARY (narration):
    We found the best boundary separating clear and turbid lakes with SVM. SVM seeks the maximum-margin separating
    surface between two classes and, via the kernel trick, can model nonlinear boundaries too. Accuracy is very high
    at 97.5% -- environmental variables separate the two regimes almost perfectly. SVM is especially powerful in
    small-to-medium, high-dimensional classification problems and is robust to overfitting thanks to margin
    maximisation. It is a strong alternative to Random Forest for ecological status classification.

====================================================================================

#65  Gradient Boosting
    file: 065_gradient_boosting_chla_nonlinear.xlsx
  >> SCENARIO (narration):
    We model how chlorophyll-a responds to multiple lake stressors, expecting
    strongly nonlinear and interactive effects rather than simple linear ones.
    For 280 observations we predict chla_ugL from trophic_status, nutrients
    TP_ugL and TN_mgL, temperature_C, water level WL_m, and macrophyte_percent.
    Gradient Boosting builds an additive ensemble of trees that captures
    curvature and threshold responses in algal biomass that linear regression
    would miss. This makes it well suited to predicting chlorophyll-a from a mix
    of categorical and continuous drivers.
  >> VARIABLE SELECTION:
    - Target: chla_ugL
    - Predictor: trophic_status
    - Predictor: TP_ugL
    - Predictor: TN_mgL
    - Predictor: temperature_C
    - Predictor: WL_m
    - Predictor: macrophyte_percent

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (trophic status classification)
    Predicted from environmental variables (sequential tree ensemble)

>> COMMENTARY (narration):
    Gradient Boosting is a powerful machine-learning method that adds trees sequentially, each new tree correcting
    the previous ones' errors. In trophic-status classification accuracy is 100% -- the model recovers the status
    perfectly from environmental variables. While Random Forest builds trees in parallel (independently), Boosting
    builds them sequentially (chasing the error); it usually gives higher accuracy but is more prone to overfitting
    and needs careful tuning. For nonlinear, interacting environmental relationships it is one of the powerful tools
    of modern ecological prediction.

====================================================================================

#66  K-Means Clustering
    file: 066_kmeans_lake_tipolojisi_4kume.xlsx
  >> SCENARIO (narration):
    We suspect that lakes fall into a few natural typologies but want to let the
    data reveal those groups rather than imposing them. For 120 lakes we cluster
    on TP_ugL, secchi_m, macrophyte_percent, and conductivity_uScm, then compare
    the discovered clusters against the known actual_typology of clear-
    macrophyte, turbid-eutrophic, saline-shallow, and deep-oligotrophic. K-Means
    Clustering partitions lakes into four compact groups by minimizing within-
    cluster variance in this environmental space. It is the right unsupervised
    method to test whether four interpretable lake types emerge from the
    measured variables.
  >> VARIABLE SELECTION:
    - Cluster variable: TP_ugL
    - Cluster variable: secchi_m
    - Cluster variable: macrophyte_percent
    - Cluster variable: conductivity_uScm
    - Validation label: actual_typology

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.78 (strong cluster structure)
    Lake typology from 4 variables (phosphorus, Secchi, macrophyte, conductivity)

>> COMMENTARY (narration):
    We split lakes into four natural groups (a typology) by their environmental features with K-Means. Silhouette =
    0.78 is very high -- the clusters are clearly separated, a real structure exists. These four groups likely
    represent distinct lake types: e.g. clear-macrophyte, eutrophic-turbid, saline, and intermediate. K-Means finds
    natural groups in unlabeled data (unsupervised learning); it is extremely practical for lake classification,
    grouping monitoring stations and building ecological typologies. The silhouette score confirms the number of
    clusters was chosen well.

====================================================================================

#67  Hierarchical Clustering
    file: 067_hierarchical_makroinvert_topluluk.xlsx
  >> SCENARIO (narration):
    We compare 25 lakes by the composition of their benthic macroinvertebrate
    communities to see which lakes are ecologically similar. Each lake carries
    abundances of Chironomidae, Oligochaeta, Ephemeroptera, Trichoptera,
    Gastropoda, Crustacea, and Coleoptera. Hierarchical Clustering builds a
    dendrogram from pairwise community distances, nesting lakes into
    successively larger groups without pre-specifying a cluster count. This is
    the natural choice for visualizing the similarity structure among lake
    communities as a tree.
  >> VARIABLE SELECTION:
    - Cluster variable: Chironomidae
    - Cluster variable: Oligochaeta
    - Cluster variable: Ephemeroptera
    - Cluster variable: Trichoptera
    - Cluster variable: Gastropoda
    - Cluster variable: Crustacea
    - Cluster variable: Coleoptera

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   Silhouette = 0.15 (weak/overlapping structure)
    Macroinvertebrate community composition

>> COMMENTARY (narration):
    We grouped macroinvertebrate communities by similarity with hierarchical clustering -- the result is a dendrogram
    (tree) and we choose the number of clusters afterward. Splitting into three clusters gives a silhouette of only
    0.15: the clusters overlap, there is no sharp separation. This is ecologically meaningful -- community
    composition often changes along a continuous gradient (downstream, eutrophication) rather than in sharp groups. A
    low silhouette means "natural grouping in the data is weak", which is itself informative. Hierarchical clustering
    is more explanatory than K-Means for seeing the nested structure of groups.

====================================================================================

#68  DBSCAN Clustering
    file: 068_dbscan_spatial_cluster_lake_koord_TP.xlsx
  >> SCENARIO (narration):
    We map 180 lakes by their geographic coordinates to find spatial hotspots of
    phosphorus pollution. Using lat and lon together with TP_ugL, we look for
    dense neighbourhoods of high-nutrient lakes while ignoring scattered
    outliers. DBSCAN Clustering groups points by spatial density, automatically
    labelling sparse, isolated lakes as noise rather than forcing them into a
    cluster. That density-based behaviour is exactly what we need to delineate
    phosphorus hotspot clusters across the landscape.
  >> VARIABLE SELECTION:
    - Coordinate: lat
    - Coordinate: lon
    - Cluster variable: TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.77   (spatial coordinates)
    Lake locations clustered by density

>> COMMENTARY (narration):
    DBSCAN clusters lakes' geographic locations by density. Unlike K-Means: we do not give the number of clusters in
    advance, and it flags low-density, isolated points as "noise" (outliers). Four dense lake groups were found
    (silhouette 0.77, strong). This shows lakes cluster into basins/regions -- e.g. the sub-basins of the Konya
    Closed Basin. DBSCAN is more appropriate than K-Means for spatial data with irregularly shaped clusters and
    outlying points (lake locations, sampling points, hot-spot detection).

====================================================================================

#69  PCA
    file: 069_pca_pan_europe_mesocosm_8cevre.xlsx
  >> SCENARIO (narration):
    Our mesocosm experiment recorded eight correlated environmental variables
    per tank, and we want to compress them into a few interpretable gradients.
    For 120 tanks we have TP_ugL, TN_mgL, chla_ugL, temperature_C, DO_mgL, pH,
    conductivity_uScm, and secchi_m. Principal Component Analysis rotates these
    intercorrelated measurements into orthogonal components, so the first axes
    capture the dominant trophic and physical gradients. This dimension
    reduction lets us visualize and interpret the main environmental structure
    across tanks.
  >> VARIABLE SELECTION:
    - Variable: TP_ugL
    - Variable: TN_mgL
    - Variable: chla_ugL
    - Variable: temperature_C
    - Variable: DO_mgL
    - Variable: pH
    - Variable: conductivity_uScm
    - Variable: secchi_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 = 46.9%   PC2 = 35.8%   (first two components total 82.7%)
    8 environmental variables

>> COMMENTARY (narration):
    We reduced eight correlated environmental variables to a few summary axes with PCA. The first two components
    explain 83% of the total variance -- so we can summarise eight variables' information in almost two axes.
    Typically PC1 represents a nutrient/eutrophication gradient (phosphorus, nitrogen, chlorophyll together) and PC2
    a physical axis (temperature, oxygen). PCA is the basic tool for visualising multivariate data, reducing
    collinearity and revealing hidden gradients. It is one of the most-used methods in limnology for summarising
    environmental monitoring data.

====================================================================================

#70  t-SNE
    file: 070_tsne_phytoplankton_topluluk_hi_dim.xlsx
  >> SCENARIO (narration):
    We want to visualize how phytoplankton communities cluster across trophic
    conditions when described by many taxa at once. For 120 samples we have
    abundances of Microcystis, Aphanizomenon, Cryptomonas, Cyclotella,
    Dinobryon, and Synedra, each labelled by trophic_status as eutrophic,
    mesotrophic, or oligotrophic. t-SNE embeds this high-dimensional community
    data into two dimensions while preserving local neighbourhood structure,
    revealing whether samples of the same trophic state group together. This
    makes t-SNE ideal for exploring community separation that linear projections
    would obscure.
  >> VARIABLE SELECTION:
    - Feature: Microcystis
    - Feature: Aphanizomenon
    - Feature: Cryptomonas
    - Feature: Cyclotella
    - Feature: Dinobryon
    - Feature: Synedra
    - Color label: trophic_status

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence (fit quality) = 0.363 -> very good
    High-dimensional phytoplankton community composition reduced to 2D

>> COMMENTARY (narration):
    t-SNE is a nonlinear method that projects high-dimensional phytoplankton community data (many species) into a
    two-dimensional plot. Unlike PCA, it focuses on preserving local neighbourhoods -- similar communities are placed
    close on the map -- ideal for seeing hidden group structure. The KL divergence is low (0.36), so the reduction
    quality is good. Important caveat: t-SNE axes and between-cluster distances are not interpretable; only the
    clustering pattern is meaningful. In community ecology it is a powerful visualisation for visually exploring
    groups of similar samples/stations.

====================================================================================

#71  MDS
    file: 071_mds_mediterranean_lake_makroinvert_benzerlik.xlsx
  >> SCENARIO (narration):
    Across thirty Mediterranean lakes, we sampled the benthic macroinvertebrate
    community to ask how similar these lakes are in their faunal composition.
    For each lake we recorded the abundances of nine taxonomic groups:
    Ephemeroptera, Plecoptera, Trichoptera, Odonata, Hemiptera, Coleoptera,
    Diptera, Mollusca and Oligochaeta. To visualise the ecological distances
    between lakes in a low-dimensional ordination, we use Multidimensional
    Scaling, which preserves the rank order of between-lake dissimilarities and
    reveals clusters of structurally similar communities. MDS is ideal here
    because it makes no strong distributional assumptions and turns a complex
    similarity matrix into an interpretable map.
  >> VARIABLE SELECTION:
    - Group: lake
    - Abundance: Ephemeroptera
    - Abundance: Plecoptera
    - Abundance: Trichoptera
    - Abundance: Odonata
    - Abundance: Hemiptera
    - Abundance: Coleoptera
    - Abundance: Diptera
    - Abundance: Mollusca
    - Abundance: Oligochaeta

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.287 (weak fit)
    Mediterranean lake macroinvertebrate similarity

>> COMMENTARY (narration):
    MDS maps the macroinvertebrate community similarity among lakes into two dimensions -- lakes with similar
    communities placed near, dissimilar ones far. Stress = 0.29 is a "weak" fit; this means community structure does
    not fully fit two dimensions, it is more complex (multidimensional) -- a common situation in ecology. Still, the
    general gradient is visible. MDS is the classic way in ecology to visualise community-similarity (e.g. Bray-
    Curtis) matrices; the stress value honestly tells how much we can trust the map.

====================================================================================

#72  UMAP
    file: 072_umap_zooplankton_topluluk.xlsx
  >> SCENARIO (narration):
    We collected 150 plankton samples spanning freshwater, mildly saline and
    highly saline waterbodies to explore how zooplankton communities reorganise
    along the salinity gradient. Each sample carries counts of five key taxa:
    daphnia, Bosmina, Cyclops, Artemia and Diaptomus. We apply UMAP to embed
    these high-dimensional community profiles into a two-dimensional manifold,
    so that samples with similar taxonomic make-up sit close together. Because
    UMAP preserves both local and broader structure, it should separate the
    fresh_macro_zoo, mild_saline and high_salt groups and expose the smooth
    transition driven by salinity tolerance.
  >> VARIABLE SELECTION:
    - Abundance: daphnia
    - Abundance: Bosmina
    - Abundance: Cyclops
    - Abundance: Artemia
    - Abundance: Diaptomus
    - Group: type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    UMAP dimension reduction (zooplankton community structure to 2D)
    n_neighbors and min_dist balance local/global structure

>> COMMENTARY (narration):
    UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure
    better and is faster. We projected zooplankton community composition into a two-dimensional map; similar
    communities cluster. The n_neighbors parameter tunes the local-global balance, and min_dist the tightness of
    clusters. UMAP is increasingly preferred for discovering hidden group structure in high-dimensional ecological/
    omics data. It is used for visualisation; the clustering pattern, not the axis values, is interpreted.

====================================================================================

#73  Cronbach's Alpha
    file: 073_cronbach_alpha_pondscape_paydas_8madde.xlsx
  >> SCENARIO (narration):
    In a multi-country survey of 180 stakeholders, we asked respondents to rate
    the importance of eight ecosystem services delivered by pondscapes. The
    eight items were pondscape_biodiversity, pondscape_water_provision,
    pondscape_climate, pondscape_aesthetic, pondscape_culture_identity,
    pondscape_recreation, pondscape_education and pondscape_conservation_value.
    Before treating these items as a single 'pondscape value' scale, we need to
    confirm they hang together reliably. Cronbach's Alpha quantifies the
    internal consistency of the eight items and tells us whether the scale is
    coherent enough to be averaged into one composite index.
  >> VARIABLE SELECTION:
    - Item: pondscape_biodiversity
    - Item: pondscape_water_provision
    - Item: pondscape_climate
    - Item: pondscape_aesthetic
    - Item: pondscape_culture_identity
    - Item: pondscape_recreation
    - Item: pondscape_education
    - Item: pondscape_conservation_value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.941 (excellent internal consistency)   Mean inter-item r = 0.667
    8-item pond/wetland value scale

>> COMMENTARY (narration):
    We tested the internal consistency of an 8-item scale measuring stakeholders' pond/wetland values with Cronbach's
    alpha. Alpha = 0.94 is excellent: the eight items consistently measure the same underlying concept (the value
    placed on wetlands), and respondents answer coherently across items. Showing a scale's reliability before
    analysing it is essential -- low alpha signals the items do not measure the same thing and the total score may be
    meaningless. Above 0.70 is acceptable, above 0.90 excellent. It is the first step of any socio-ecological study
    that develops a scale.

====================================================================================

#74  Likert Analysis
    file: 074_likert_5li_paydas_10madde.xlsx
  >> SCENARIO (narration):
    We surveyed 220 stakeholders across several European countries about their
    attitudes toward lake and pond management, using ten five-point Likert
    statements from item_1 to item_10. We want to understand the distribution of
    agreement and disagreement for each statement and whether responses cluster
    toward consensus or polarisation. Likert Analysis summarises the ordinal
    response patterns item by item, producing diverging stacked frequencies that
    make agreement structure immediately readable. This approach respects the
    ordinal nature of the data rather than forcing it into a mean, which is
    exactly what these perception items require.
  >> VARIABLE SELECTION:
    - Item: item_1
    - Item: item_2
    - Item: item_3
    - Item: item_4
    - Item: item_5
    - Item: item_6
    - Item: item_7
    - Item: item_8
    - Item: item_9
    - Item: item_10
    - Group: country
    - Group: age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.048 (UNACCEPTABLE)   10 items   n = 220   item mean ~ 3.0
    Inter-item correlations very low

>> COMMENTARY (narration):
    This 5-point Likert scale gave a striking result: Cronbach's alpha is only 0.05 -- unacceptably low. This means
    the items do NOT measure a common concept; it is as if 10 unrelated questions were asked. Summing such items into
    a single score would be wrong. This is instructive: a high mean or pretty charts do not make a scale valid;
    reliability (alpha) must be checked first. A low alpha says the scale must be redesigned, the items revised, or
    split into sub-dimensions. Likert analysis exists precisely for this diagnosis -- it exposes early whether a
    scale is sound.

====================================================================================

#75  EFA
    file: 075_efa_paydas_3faktor_12madde.xlsx
  >> SCENARIO (narration):
    From a questionnaire of 200 stakeholders we measured twelve items intended
    to capture how people value pondscapes, and we suspect these items reflect a
    smaller number of latent dimensions. The items span three conceptual blocks:
    extra_service_1 to extra_service_4, cultural_5 to cultural_8, and economic_9
    to economic_12. We run Exploratory Factor Analysis to uncover the underlying
    factor structure and see whether the items load cleanly onto the expected
    service, cultural and economic factors. EFA is appropriate because we are
    exploring, not confirming, the dimensionality of a newly assembled value
    scale.
  >> VARIABLE SELECTION:
    - Item: extra_service_1
    - Item: extra_service_2
    - Item: extra_service_3
    - Item: extra_service_4
    - Item: cultural_5
    - Item: cultural_6
    - Item: cultural_7
    - Item: cultural_8
    - Item: economic_9
    - Item: economic_10
    - Item: economic_11
    - Item: economic_12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO (overall) = 0.827 (very good)   -> factor analysis appropriate
    12 items, expected 3 factors (extra service / cultural / economic)

>> COMMENTARY (narration):
    We explored how many hidden dimensions (factors) underlie twelve survey items with Exploratory Factor Analysis.
    KMO = 0.83 is "very good": the data are suitable for factor analysis, with enough shared variance among items.
    The items group as expected into three factors -- extra services, cultural values, economic values. EFA reduces
    many items to a few meaningful dimensions, both simplifying the scale and revealing the underlying structures. It
    is a core tool of scale development and psychometric/socio-ecological research; KMO and Bartlett's test are
    reported as the pre-check.

====================================================================================

#76  ICC
    file: 076_icc_3lab_TP_measurements.xlsx
  >> SCENARIO (narration):
    Three laboratories independently measured total phosphorus in water samples
    from the same 30 lakes, recorded as lab1_TP_ugL, lab2_TP_ugL and
    lab3_TP_ugL. Before pooling these measurements into a regional monitoring
    database, we must know whether the labs agree well enough to be treated as
    interchangeable. We compute the Intraclass Correlation Coefficient to
    quantify the absolute agreement among the three labs on the same lakes. ICC
    is the right tool because it captures both correlation and systematic bias
    between raters measuring a continuous quantity.
  >> VARIABLE SELECTION:
    - Rater: lab1_TP_ugL
    - Rater: lab2_TP_ugL
    - Rater: lab3_TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.935   ICC(C,1) = 0.942   (Excellent)   3 laboratories
    95% CI ~ (0.88 , 0.97)

>> COMMENTARY (narration):
    We assessed how consistent three laboratories are in measuring the phosphorus value of the same water samples
    with ICC. ICC = 0.94 is "excellent": inter-lab measurement agreement is very high -- whichever lab measures, the
    result is similar. This is critical in multi-centre/multi-observer studies: if labs were inconsistent, we could
    not tell whether observed differences are real or measurement error. ICC is the standard for measuring observer/
    instrument/lab reliability on continuous measurements (Cohen's Kappa is for categorical data, ICC for continuous).

====================================================================================

#77  CFA
    file: 077_cfa_sem_pondscape_yer_identity.xlsx
  >> SCENARIO (narration):
    Building on theory, we hypothesise that stakeholders' attachment to
    pondscapes is composed of two distinct latent constructs: place identity,
    measured by identity_1 to identity_4, and place dependence, measured by
    dependence_1 to dependence_4. With responses from 200 participants, we want
    to test whether the data actually conform to this two-factor measurement
    model. Confirmatory Factor Analysis lets us specify the expected loading of
    each indicator on its construct and assess model fit directly. CFA suits
    this stage because the structure is theory-driven and we are confirming, not
    exploring, the two-dimensional attachment model.
  >> VARIABLE SELECTION:
    - Identity indicator: identity_1
    - Identity indicator: identity_2
    - Identity indicator: identity_3
    - Identity indicator: identity_4
    - Dependence indicator: dependence_1
    - Dependence indicator: dependence_2
    - Dependence indicator: dependence_3
    - Dependence indicator: dependence_4
    - Group: country

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.998   RMSEA = 0.021   (excellent fit)
    2 factors: place identity + place dependence

>> COMMENTARY (narration):
    While EFA explores hidden structure, CFA tests a pre-specified theory: here we tested whether two distinct factors
    -- "place identity" and "place dependence" -- fit the data. The fit is excellent (CFI = 0.998, RMSEA = 0.021): the
    items really do load onto these two dimensions, the hypothesised scale structure is confirmed. CFA is part of
    structural equation modelling (SEM) and is the strongest way to prove scale validity. In environmental-psychology
    studies measuring people's emotional/functional bonds to wetlands, this is how we demonstrate scale validity.

====================================================================================

#78  Survey Means
    file: 078_survey_means_chla_regional.xlsx
  >> SCENARIO (narration):
    Using a stratified, clustered monitoring design covering 220 sites, we want
    to estimate the average chlorophyll-a concentration across Turkey's lake
    regions while properly accounting for the sampling structure. The design
    includes regional strata (stratum_region), sampling clusters (cluster_id)
    and survey weights, with chla_ugL as the primary response and TP_ugL as a
    secondary nutrient indicator. Survey Means computes design-based estimates
    of the population mean with correct standard errors that reflect the
    weighting and clustering. This method is essential because ignoring the
    complex design would underestimate the uncertainty of our chlorophyll-a
    estimates.
  >> VARIABLE SELECTION:
    - Stratum: stratum_region
    - Cluster: cluster_id
    - Weight: weight
    - Mean variable: chla_ugL
    - Mean variable: TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Chlorophyll weighted mean = 26.77   SE = 8.42   95% CI (10.02 , 43.53)   CV 31.5%
    Stratified + clustered + weighted design (Taylor SE)

>> COMMENTARY (narration):
    In a regional monitoring survey, lakes were not chosen at random but with a stratified, clustered and weighted
    design. A simple mean would be wrong for such data; complex-sample analysis gives the correct mean and standard
    error (Taylor linearization) by accounting for the design weights and clustering. The weighted mean chlorophyll
    is 26.8, with a wide 95% CI (the design effect increases uncertainty). National/regional environmental monitoring
    programs (water-quality inventories, national health surveys) are always designed this way; survey methods are
    essential for unbiased estimation.

====================================================================================

#79  Survey Frequencies
    file: 079_survey_frequency_trophic_status_distribution.xlsx
  >> SCENARIO (narration):
    From a stratified, clustered survey of 240 lake sites we want to estimate
    how lakes are distributed across trophic categories at the national scale.
    Each site has a regional stratum (stratum_region), a sampling cluster
    (cluster_id), a survey weight, and a classified trophic_status of hyper, eu,
    meso or oligo. Survey Frequencies produces design-based estimates of the
    proportion of lakes in each trophic class, with confidence intervals that
    honour the weighting and clustering. This is the appropriate approach for
    turning weighted categorical observations into valid population-level
    prevalence estimates.
  >> VARIABLE SELECTION:
    - Stratum: stratum_region
    - Cluster: cluster_id
    - Weight: weight
    - Category: trophic_status

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Eutrophic (eu) proportion = 0.320   SE = 0.059   95% CI (0.215 , 0.447)
    Stratified + clustered + weighted design

>> COMMENTARY (narration):
    We estimated the proportion of trophic-status categories (oligo/meso/eu/hyper) among lakes, accounting for the
    complex sample design. The estimated proportion of eutrophic lakes is 32% (SE 5.9%). A simple percentage would
    give the wrong standard error by ignoring stratification and clustering; survey frequency analysis uses the
    weights and design to produce a confidence interval reflecting the true uncertainty. It is the correct way to
    estimate category proportions in a population (eutrophic lake share, disease prevalence, urbanisation rate) in a
    design-consistent manner.

====================================================================================

#80  Survey Totals
    file: 080_survey_total_wetland_area_total.xlsx
  >> SCENARIO (narration):
    Across 180 sampled basins drawn under a stratified, clustered design, we aim
    to estimate the total wetland area for the entire study region rather than
    just an average. Each basin record contains a stratum, a cluster_id, a
    survey weight, and the measured wetland_area_ha. Survey Totals expands the
    weighted sample to a design-based estimate of the grand total wetland area,
    complete with a standard error that reflects the survey structure. We use
    this method because management and conservation targets are expressed as
    absolute hectares, requiring a population total rather than a mean.
  >> VARIABLE SELECTION:
    - Stratum: stratum
    - Cluster: cluster_id
    - Weight: weight
    - Total variable: wetland_area_ha

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total wetland area = 86,727.73 ha   SE = 10,377.09   95% CI (66,249 , 107,206)
    Stratified design (Taylor SE)

>> COMMENTARY (narration):
    We estimated the total wetland area in a region using the sampling design and weights: about 86,700 hectares (95%
    CI 66,000 - 107,000). Scaling from the sampled basins up to the whole population requires correct weighting --
    each sampled unit carries weight proportional to the population fraction it represents. Survey total analysis does
    exactly this and expresses uncertainty with the standard error. It is the official-statistics method used to
    estimate population totals (total forest area, total population, total production) from a sample.

====================================================================================

#81  Survey Regression
    file: 081_survey_regression_chla_TP.xlsx
  >> SCENARIO (narration):
    In our national lake monitoring program, sampling was not random but drawn
    from trophic strata with unequal probabilities, so each lake carries a
    survey weight and sits within a sampling cluster. We want to know how total
    phosphorus, TP_ugL, together with water temperature_C, drives chlorophyll-a,
    chla_ugL, while respecting that design. Survey-weighted regression lets us
    estimate these nutrient effects on algal biomass with correct standard
    errors that account for the stratum and cluster_id structure. This gives us
    population-level inference about eutrophication rather than biased estimates
    from treating the sample as simple random.
  >> VARIABLE SELECTION:
    - Stratum: stratum
    - Cluster: cluster_id
    - Weight: weight
    - Predictor: TP_ugL
    - Predictor: temperature_C
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.908   TP coefficient = 0.441   SE = 0.009   p < .001   95% CI (0.424 , 0.458)
    Stratified + clustered + weighted regression

>> COMMENTARY (narration):
    We modelled the chlorophyll-phosphorus relationship with a regression accounting for the complex sample design
    (strata, clusters, weights). Phosphorus's effect is strong and significant (coefficient 0.44, p < .001), and the
    model explains 91% of the variance. Design-based standard errors are more accurate than the (often too
    optimistic) SEs from simple OLS. When estimating relationships from national/regional monitoring data (where the
    sample is not random), survey regression is necessary for valid inference; otherwise confidence intervals and
    p-values would mislead.

====================================================================================

#82  Survey Logistic Regression
    file: 082_survey_logistic_turbid.xlsx
  >> SCENARIO (narration):
    From the same stratified, clustered lake survey we ask a yes/no question: is
    a given station turbid, coded in turbid_1_0. We hypothesize that rising
    total phosphorus, TP_ugL, increases the odds of a lake tipping into a turbid
    state. Survey logistic regression models this binary outcome while
    incorporating the survey weight, the stratum, and the cluster_id so our odds
    ratios are valid for the whole lake population. This tells us how strongly
    nutrient enrichment predicts turbidity across the monitored region.
  >> VARIABLE SELECTION:
    - Stratum: stratum
    - Cluster: cluster_id
    - Weight: weight
    - Predictor: TP_ugL
    - Outcome: turbid_1_0

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TP coefficient = 0.0237   OR = 1.024   p < .001   95% CI (1.014 , 1.034)
    Stratified + clustered + weighted logistic

>> COMMENTARY (narration):
    Predicting the probability that a lake is turbid from phosphorus, we handled both the binary outcome (logistic)
    and the complex sample design together. The odds ratio is 1.024 (p < .001): for each 1-unit rise in phosphorus,
    the odds of being turbid rise about 2.4% -- cumulatively a large difference. Because the design weights and
    clustering are accounted for, the OR and confidence interval are design-consistent and valid. Survey logistic
    regression is the correct way to model binary outcomes (disease yes/no, turbid/clear, yes/no) from complex
    surveys.

====================================================================================

#83  GAM
    file: 083_gam_smooth_temperature_TP_chla.xlsx
  >> SCENARIO (narration):
    Algal response to nutrients and warming is rarely a straight line, so a
    simple linear model may miss thresholds and saturation in lake productivity.
    Here we model chlorophyll-a, chla_ugL, as a smooth function of both total
    phosphorus, TP_ugL, and water temperature_C. A generalized additive model
    fits flexible spline smooths that reveal where chla rises steeply with TP
    and where temperature amplifies blooms, without forcing a predefined curve.
    This lets us visualize the true shape of the nutrient-temperature-algae
    relationship.
  >> VARIABLE SELECTION:
    - Smooth predictor: temperature_C
    - Smooth predictor: TP_ugL
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (explained) = 0.867   McFadden adj = 0.493
    Chlorophyll ~ s(temperature) + s(phosphorus)  (smooth functions)

>> COMMENTARY (narration):
    GAM is a flexible generalisation of linear regression: it models a variable's effect with "smooth curves"
    (splines) rather than a straight line. Chlorophyll's relationship with temperature and phosphorus is often
    nonlinear -- it accelerates or saturates beyond certain thresholds. GAM learns this curved pattern from the data
    without assuming its shape in advance; the model explains 87% of the variance. It preserves interpretability
    while adding flexibility: we can read from the curve where each variable accelerates. For nonlinear but
    interpretable relationships in ecology, it is an ideal method.

====================================================================================

#84  Discriminant Analysis (LDA/QDA)
    file: 084_diskriminant_3durum_lake_classification.xlsx
  >> SCENARIO (narration):
    Lakes in our survey fall into three ecological regimes recorded in
    lake_status: clear, medium, and turbid. We want to know which limnological
    measurements best separate these states and to build a rule that classifies
    a new lake. Discriminant analysis uses total phosphorus, TP_ugL, Secchi
    depth secchi_m, macrophyte cover macrophyte_percent, and chlorophyll-a
    chla_ugL to find the combinations that maximally distinguish the three
    regimes. LDA or QDA then gives us an interpretable classifier and shows
    which variables drive the clear-to-turbid gradient.
  >> VARIABLE SELECTION:
    - Group: lake_status
    - Predictor: TP_ugL
    - Predictor: secchi_m
    - Predictor: macrophyte_percent
    - Predictor: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (3 lake-status classification)
    Separation from 4 environmental variables

>> COMMENTARY (narration):
    Discriminant analysis finds the linear combinations of variables that best separate the groups (three lake
    states) -- a bit like a classification-focused PCA. Accuracy is 100%: four environmental variables separate the
    three states perfectly. Discriminant analysis both classifies and answers "which variables matter most for the
    separation" (discriminant function loadings). It resembles logistic regression but is the classic choice for
    multiple groups when variables are normally distributed. It is used in ecological status/type classification and
    for assigning new observations to existing groups.

====================================================================================

#85  Conditional Logit
    file: 085_conditional_logit_lake_visit_choice.xlsx
  >> SCENARIO (narration):
    To understand recreational pressure on lakes, we ran a choice experiment
    where each subject was offered several lake alternatives and picked one,
    marked in chosen. Each alternative is described by its distance_km from the
    visitor and its water_quality_1_5 rating. Conditional logit models the
    probability that an alternative is chosen as a function of these
    alternative-specific attributes within each choice set. This reveals how
    much visitors trade off travel distance against water quality when selecting
    a lake.
  >> VARIABLE SELECTION:
    - Alternative: alternative
    - Attribute: distance_km
    - Attribute: water_quality_1_5
    - Choice: chosen

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (McFadden) = 0.324
    Lake-visit choice ~ distance (km) + water quality (1-5)

>> COMMENTARY (narration):
    We examined what people look at when choosing which lake to visit with a conditional logit model -- each person
    selects one from a choice set. McFadden pseudo R2 = 0.32 is quite good for choice models (0.2-0.4 is considered
    strong). The model shows visitors prefer nearby lakes with high water quality. This method underlies "discrete
    choice" analysis in transport, marketing and environmental economics (recreation demand, ecosystem-service
    valuation). It is powerful for modelling wetland recreation value and visitor behaviour.

====================================================================================

#86  Kaplan-Meier
    file: 086_kaplan_meier_daphnia_DVM_kairomone.xlsx
  >> SCENARIO (narration):
    In a predation-cue experiment, individual Daphnia were exposed either to
    fish_kairomone or to control water, and we tracked how long each survived
    during diel vertical migration trials. The duration_hour records observed
    time and death_event marks whether the animal actually died or was censored.
    Kaplan-Meier estimates and compares the survival curves of the two treatment
    groups over time. This shows whether the chemical presence of a predator
    shortens Daphnia survival.
  >> VARIABLE SELECTION:
    - Group: treatment
    - Time: duration_hour
    - Event: death_event

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 72 (90%)   Median survival = 16.61 hours
    Log-rank: chi-square(1) = 55.83, p < .001  (groups differ significantly)

>> COMMENTARY (narration):
    We analysed Daphnia (water flea) survival time -- with and without kairomone (fish scent) present -- using
    Kaplan-Meier. Median survival is 16.6 hours, and the log-rank test shows a highly significant difference between
    groups (chi-square = 55.83, p < .001): the fish scent (predator cue) clearly changes survival. Kaplan-Meier
    analyses "time-to-event" data, correctly using censored (event-not-yet-occurred) observations too. It is the
    fundamental visual and test for "time-event" data such as survival/death, pond drying time and macrophyte loss
    time.

====================================================================================

#87  Cox Regression
    file: 087_cox_PH_lake_regime_transition.xlsx
  >> SCENARIO (narration):
    We followed a set of lakes over time to see how long each remained stable
    before a regime_transition from clear to turbid occurred, with follow-up
    recorded in duration_month. We suspect that total phosphorus, TP_ugL,
    macrophyte cover macrophyte_percent, and conductivity_uScm govern the hazard
    of flipping. Cox proportional-hazards regression estimates how each
    covariate raises or lowers the instantaneous risk of transition without
    assuming a baseline shape. This identifies which drivers accelerate the loss
    of the clear-water state.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: regime_transition
    - Predictor: TP_ugL
    - Predictor: macrophyte_percent
    - Predictor: conductivity_uScm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 94 (94%)   Concordance = 0.753
    TP: HR = 1.014 p < .001   |   macrophyte: HR = 0.970 p < .001 (protective)   |   conductivity: HR = 1.001 p < .001

>> COMMENTARY (narration):
    We modelled the risk of a lake's regime shift (clear to turbid) from environmental variables with Cox regression.
    The model discriminates well (concordance 0.75). The results are clear ecologically: phosphorus increases risk
    (HR > 1), while macrophytes are protective (HR = 0.97 < 1 -- more aquatic plants, lower shift risk). Cox
    regression is the gold-standard survival model relating "time-to-event" to several predictors; it gives hazard
    ratios (HR) -- "how many times faster regime collapse occurs as phosphorus rises". It is very valuable for
    understanding lake stability and transition risk.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT)
    file: 088_parametrik_AFT_weibull_macrophyte_loss.xlsx
  >> SCENARIO (narration):
    Here we model the time until submerged macrophytes are lost from a lake,
    recorded as duration_month_macrophyte_loss with loss_event indicating
    whether collapse occurred. We want a model that directly describes how total
    phosphorus, TP_ugL, conductivity_uScm, and fish abundance fish_cpue stretch
    or shorten the time to that collapse. A parametric Weibull accelerated
    failure time model expresses these covariates as multiplicative effects on
    survival time, giving intuitive time-ratio interpretations. This quantifies
    how nutrient and fish pressure hasten macrophyte loss.
  >> VARIABLE SELECTION:
    - Predictor: TP_ugL
    - Predictor: conductivity_uScm
    - Predictor: fish_cpue
    - Time: duration_month_macrophyte_loss
    - Event: loss_event

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution = Weibull   Events = 85 (94.4%)   Median survival = 40.02 months
    Macrophyte loss time ~ phosphorus + conductivity + fish (CPUE)

>> COMMENTARY (narration):
    We analysed the time until macrophyte (aquatic plant) cover is lost with a parametric Weibull survival model.
    While Cox is semi-parametric, Weibull models the whole survival curve with a specific mathematical distribution
    -- if the data fit that distribution it gives more efficient estimates and allows extrapolation. Median survival
    is 40 months: in a typical lake, macrophyte loss occurs on average over this period. The AFT (accelerated failure
    time) interpretation is also intuitive: how many times a variable lengthens/shortens the time. When the time
    distribution is known, parametric models are powerful.

====================================================================================

#89  Competing Risks
    file: 089_competing_risks_lake_last_status.xlsx
  >> SCENARIO (narration):
    Lakes we monitored over years ended in one of several mutually exclusive
    fates captured in result_0sansurlu_1kuruma_2tuzlanma_3restorasyon: still
    censored, dried out, salinized, or restored. Because these outcomes compete
    with one another, a standard survival model would overstate the risk of any
    single fate. Competing-risks analysis uses duration_year together with total
    phosphorus, TP_ugL, and conductivity_uScm to estimate the cumulative
    incidence of each ending separately. This tells us how nutrient and salinity
    pressure shift a lake toward drying, salinization, or recovery.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Predictor: TP_ugL
    - Predictor: conductivity_uScm
    - Event: result_0sansurlu_1kuruma_2tuzlanma_3restorasyon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 44   |   Drying (1) = 27   |   Salinization (2) = 25   |   Restoration (3) = 24
    Separate cumulative incidence (CIF) and risk for each event type

>> COMMENTARY (narration):
    A lake can meet more than one "final state": drying, salinization, or (positively) restoration -- and these are
    mutually exclusive (once one happens the others cannot). Classic survival handles this incorrectly; competing-
    risks analysis produces a separate cumulative incidence for each event type. In the data, 27 lakes dried, 25
    salinized, 24 were restored, and 44 are still at risk. This framework correctly answers questions like "which
    final state's risk does phosphorus increase". It is essential when competing events exist -- in medicine
    (different causes of death) and ecology (different degradation pathways).

====================================================================================

#90  Time-Dependent Cox
    file: 090_time_dep_cox_time_variable_TP_macr.xlsx
  >> SCENARIO (narration):
    Because total phosphorus and macrophyte cover in a lake change while we
    observe it, treating them as fixed baseline values would be misleading. Our
    data are arranged in start-stop intervals per lake, with TP_ugL and
    macrophyte_percent updated across intervals and event marking when the
    transition happens. A time-dependent Cox model lets these covariates vary
    over follow-up so the hazard reflects each lake's current condition. This
    captures how evolving nutrient and vegetation states drive the timing of
    regime change.
  >> VARIABLE SELECTION:
    - Start: start
    - Stop: stop
    - Time-varying predictor: TP_ugL
    - Time-varying predictor: macrophyte_percent
    - Event: event

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TP: beta = +0.033   HR = 1.033   95% CI (1.011 , 1.057)   z = 2.88   p = .004
    Time-varying phosphorus/macrophyte covariates

>> COMMENTARY (narration):
    Sometimes predictors change over time: a lake's phosphorus and macrophytes fluctuate over the years. Standard Cox
    assumes them constant; time-dependent Cox uses the current value in each time interval (start-stop format).
    Phosphorus is significant (HR = 1.033, p = .004): at any moment, high phosphorus raises the regime-collapse risk
    at that moment. This is the correct method when "the current, continuously monitored condition matters, not just
    the baseline value". It is the standard for modelling time-varying exposures (pollution, treatment, climate) in
    long-term monitoring.

====================================================================================

#91  Survey-Weighted Cox
    file: 091_survey_phreg_complex_survey_survival.xlsx
  >> SCENARIO (narration):
    We monitored 100 lakes drawn from a complex survey design, where stations
    were nested within clusters and grouped into three sampling strata to track
    how long each lake remained below a critical water-quality threshold. Using
    duration_year as the follow-up time and event as the indicator of crossing
    into eutrophic status, we ask whether total phosphorus TP_ugL accelerates
    that transition. Because the lakes were selected with unequal probabilities,
    each carries a survey weight, and observations are organized by stratum and
    cluster_id. A survey-weighted Cox model is the right tool here, since it
    estimates the hazard of regime change for TP_ugL while properly accounting
    for the stratified, clustered, weighted design.
  >> VARIABLE SELECTION:
    - Stratum: stratum
    - Cluster: cluster_id
    - Weight: weight
    - Time: duration_year
    - Event: event
    - Covariate: TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 64   Concordance = 0.708   AIC (partial) = 1060.26
    Weighted/clustered survival (phosphorus covariate)

>> COMMENTARY (narration):
    We combined survival analysis with a complex sample design: observations come from a population sampled with
    weights and clusters. Survey-PHREG runs the Cox model with design weights and cluster-robust standard errors, so
    the hazard ratios (HR) and confidence intervals generalise to the population. Concordance 0.71 means the model
    discriminates reasonably. When national health surveys or scaled environmental monitoring programs have "time-to-
    event" data, ignoring the design biases the estimates; survey-PHREG provides valid inference.

====================================================================================

#92  Interval-Censored Survival
    file: 092_interval_censored_lake_regime_interval.xlsx
  >> SCENARIO (narration):
    For 80 lakes we could not pinpoint the exact month each one shifted to a
    degraded trophic regime, because sampling happened only at intervals, so for
    every lake we know only a lower_bound_month and an upper_bound_month
    bracketing the true transition time. We want to understand how total
    phosphorus TP_ugL relates to the timing of this regime shift despite that
    uncertainty. Since the event time is known only within an interval rather
    than exactly, ordinary survival methods would bias the estimates. Interval-
    censored survival analysis is appropriate, modeling the transition time
    between lower_bound_month and upper_bound_month as a function of TP_ugL.
  >> VARIABLE SELECTION:
    - Lower bound: lower_bound_month
    - Upper bound: upper_bound_month
    - Covariate: TP_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 58   Median survival = 107.0 months   (Turnbull algorithm)
    Event time known within [lower, upper] interval

>> COMMENTARY (narration):
    In field monitoring we rarely know the exact time of an event: we visit a lake once a year, find it clear one day
    and turbid at the next visit -- the shift happened somewhere between those two dates. Interval-censored survival
    (Turnbull algorithm) handles exactly this uncertainty, avoiding fixing the event to an arbitrary date. Median
    survival is 107 months. By the nature of periodic monitoring (annual sampling, periodic diagnosis), event times
    are always within an interval; forcing this data to a single point biases the result. The interval-censored
    method matches the reality of monitoring data.

====================================================================================

#93  Frailty Cox
    file: 093_frailty_cox_basin_cluster_frailty.xlsx
  >> SCENARIO (narration):
    We followed 150 lakes grouped into shared drainage basins to study how long
    each lake stayed in good ecological condition before an event occurred.
    Lakes within the same basin_cluster experience correlated environmental
    pressures, so their failure times are not independent. We test whether total
    phosphorus TP_ugL and conductivity_uScm shorten the time to event recorded
    in duration_year, while a frailty term captures the unobserved basin-level
    susceptibility. A frailty Cox model fits this structure, adding a random
    effect for basin_cluster on top of the hazard driven by TP_ugL and
    conductivity_uScm.
  >> VARIABLE SELECTION:
    - Cluster: basin_cluster
    - Time: duration_year
    - Event: event
    - Covariate: TP_ugL
    - Covariate: conductivity_uScm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 148   Concordance = 0.763   TP: HR = 1.019 p < .001
    Basin-clustered shared frailty (random effect)

>> COMMENTARY (narration):
    Lakes are clustered within basins; lakes in the same basin share unmeasured common factors (geology, climate,
    land use) and carry similar risk. Frailty Cox adds a shared "frailty" (random effect) for each basin to model
    this clustering -- the mixed-model version of survival analysis. Phosphorus is again significant (HR = 1.019,
    p < .001), concordance 0.76. Accounting for basin-level hidden differences gives both correct standard errors and
    information on "how much heterogeneity exists between clusters". It is the required method for clustered survival
    data (hospitals, basins, families).

====================================================================================

#94  Time Series Description
    file: 094_time_series_eymir_20yil_monthly.xlsx
  >> SCENARIO (narration):
    We assembled a 20-year monthly record from Lake Eymir, with 240 observations
    indexed by date, year, and month, capturing chlorophyll-a chla_ugL and water
    level WL_m. Before any modeling, we want a clear descriptive picture of how
    chla_ugL and WL_m behave over time, including their central tendency,
    spread, and seasonal rhythm. A time series description is the natural first
    step, summarizing the chla_ugL and WL_m series across the date axis to
    reveal the basic temporal structure.
  >> VARIABLE SELECTION:
    - Date: date
    - Year: year
    - Month: month
    - Series: chla_ugL
    - Series: WL_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    240 observations (monthly)   ADF p = 0.415 -> NOT stationary   Trend = decreasing   Seasonality detected

>> COMMENTARY (narration):
    We examined the 20-year monthly chlorophyll series of Lake Eymir. The ADF test does not find the series
    stationary (p = 0.42) -- the mean/variance change over time, there is a trend (decreasing, likely a restoration
    effect) and strong seasonality is present. This pre-diagnosis is critical: most time-series models (like ARIMA)
    require stationarity, so differencing/transformation may be needed first. Time-series analysis reveals the
    structure of long monitoring data -- trend, seasonality, stationarity, autocorrelation -- and tells which
    modelling steps are required. It is the foundation of water-quality monitoring.

====================================================================================

#95  STL Decomposition
    file: 095_stl_mogan_monthly_chla_decomp.xlsx
  >> SCENARIO (narration):
    Using a 180-month chlorophyll-a record from Lake Mogan, indexed by date with
    year and month, we want to separate the algal signal into its underlying
    components. The chla_ugL series clearly contains a long-term tendency plus a
    repeating annual cycle, and we need to isolate these from short-term noise.
    STL decomposition is well suited here, splitting chla_ugL into trend,
    seasonal, and remainder parts so we can interpret the seasonal phytoplankton
    dynamics distinctly from the overall direction.
  >> VARIABLE SELECTION:
    - Date: date
    - Year: year
    - Month: month
    - Series: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Period (seasonal length) = 12 (monthly -> annual cycle)
    Series decomposed into trend + seasonal + residual components

>> COMMENTARY (narration):
    STL (Seasonal-Trend decomposition using Loess) splits Lake Mogan's monthly chlorophyll series into three
    components: long-term trend, a recurring seasonal cycle (12-month) and the remaining residual (noise/events).
    This decomposition clarifies "is chlorophyll increasing, or just peaking every summer" -- once the seasonality is
    removed we can see the true underlying trend. STL is the most intuitive way to interpret seasonal environmental
    series (chlorophyll, temperature, water level), separating trend from seasonal effect so each can be assessed
    on its own.

====================================================================================

#96  ARIMA
    file: 096_arima_beysehir_monthly_chla_forecast.xlsx
  >> SCENARIO (narration):
    We have 180 consecutive monthly chlorophyll-a measurements from Lake
    Beysehir, indexed by date, and we want to forecast future algal biomass. The
    chla_ugL series shows autocorrelation and seasonal behavior that we can
    exploit to project the coming months. An ARIMA model is appropriate, fitting
    the autoregressive and moving-average structure of chla_ugL over date to
    generate short-term forecasts of phytoplankton conditions.
  >> VARIABLE SELECTION:
    - Date: date
    - Series: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model fitted   AIC = 1136.78
    Beysehir monthly chlorophyll -- future forecast

>> COMMENTARY (narration):
    We fitted an ARIMA model to Lake Beysehir's monthly chlorophyll series and forecast the future. ARIMA combines
    the series' own past values (AR), past errors (MA) and differencing (I, to make it stationary). With AIC = 1137
    the best model was selected. The forecasts rest on the observed trend and autocorrelation structure and come with
    an uncertainty band. Time-series forecasting matters for water-quality management: anticipating future
    eutrophication risk enables early warning and resource planning. ARIMA is the classic forecasting method for
    environmental series with seasonality and trend.

====================================================================================

#97  Exponential Smoothing
    file: 097_ets_holt_winters_chla_mevsimsel.xlsx
  >> SCENARIO (narration):
    From a 144-month chlorophyll-a record indexed by date with year and month,
    we aim to forecast seasonal algal blooms in a lake. The chla_ugL series
    displays both a gradual level change and a strong recurring yearly pattern.
    Holt-Winters exponential smoothing fits this case, modeling the level,
    trend, and seasonality of chla_ugL to produce smooth seasonal forecasts.
  >> VARIABLE SELECTION:
    - Date: date
    - Year: year
    - Month: month
    - Series: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method = Holt-Winters (additive seasonal, period 12)   AIC = 420.92
    Seasonal chlorophyll forecast

>> COMMENTARY (narration):
    Holt-Winters exponential smoothing estimates the level + trend + seasonality components in a weighted way (giving
    more weight to the recent past). It captured the seasonal pattern of chlorophyll (AIC = 421, lower than ARIMA --
    a better fit for this series). ETS/Holt-Winters is often very successful for series with strong seasonality and
    is more intuitive than ARIMA. In environmental monitoring forecasts it is an easy-to-build, interpretable
    alternative, especially preferred for regularly seasonally fluctuating variables (chlorophyll, temperature,
    demand).

====================================================================================

#98  Mann-Kendall Trend
    file: 098_mann_kendall_beysehir_WL_40yil_trend.xlsx
  >> SCENARIO (narration):
    We compiled 40 years of annual data for Lake Beysehir, with one record per
    year, tracking water level WL_m alongside precipitation_mm and
    temperature_C. We want to know whether the lake's water level shows a
    significant monotonic long-term trend rather than random fluctuation.
    Because the data are annual and may not meet normality assumptions, a non-
    parametric Mann-Kendall trend test is appropriate to assess directional
    change in WL_m over the years, with precipitation_mm and temperature_C as
    accompanying climatic context.
  >> VARIABLE SELECTION:
    - Year: year
    - Series: WL_m
    - Series: precipitation_mm
    - Series: temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   Trend = DECREASING   Sen's slope = -0.186 unit/observation
    Beysehir water level 40-year trend

>> COMMENTARY (narration):
    We tested whether Lake Beysehir's 40-year water level shows a significant trend with the Mann-Kendall test -- a
    nonparametric, outlier- and distribution-robust trend test. The result is highly significant (p < .001) and the
    direction is decreasing: the Sen's slope is -0.186, i.e. water level falls consistently at each measurement. This
    statistically documents the drying crisis of Anatolia's closed basins. Sen's slope gives the trend magnitude
    without being affected by outliers. For long-term trend detection in climate and hydrology series, Mann-Kendall
    + Sen's is the gold standard.

====================================================================================

#99  Anomaly Detection
    file: 099_anomali_tespiti_chla_monthly_outlier.xlsx
  >> SCENARIO (narration):
    We examine a 240-month chlorophyll-a series, indexed by date, from a lake
    under intermittent stress. Within the otherwise regular chla_ugL signal we
    suspect occasional extreme bloom events or sensor irregularities that
    deviate sharply from expected behavior. Anomaly detection is the right
    approach, scanning the chla_ugL series along date to flag observations that
    depart from the normal temporal pattern.
  >> VARIABLE SELECTION:
    - Date: date
    - Series: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Anomaly rate = 2.08%   (IQR-based outlier detection)
    Extraordinary values in the monthly chlorophyll series

>> COMMENTARY (narration):
    We detected the extraordinary values (anomalies) in the monthly chlorophyll series -- about 2% of observations
    lie outside the normal fluctuation. These anomalies usually point to real ecological events: a sudden algal bloom,
    a pollution event, an extreme heatwave, or measurement error. Anomaly detection automatically answers "when did
    something go wrong" in continuous monitoring data and is the basis of early-warning systems. In water-quality
    monitoring it is extremely practical for picking out the few critical events that require attention from among
    thousands of measurements.

====================================================================================

#100  Variance Components (VARCOMP)
    file: 100_varcomp_basin_lake_station_chla.xlsx
  >> SCENARIO (narration):
    We collected 240 chlorophyll-a measurements hierarchically nested across
    four basins, with lake_id within basins, station_id within lakes, and
    repeated samples per station. We want to know at which spatial level the
    variability in chla_ugL is concentrated, whether differences arise mostly
    between basins, between lakes, between stations, or within repeated
    sampling. Variance components analysis (VARCOMP) is appropriate,
    partitioning the total variance of chla_ugL into the contributions of basin,
    lake_id, station_id, and residual repeat measurement.
  >> VARIABLE SELECTION:
    - Factor: basin
    - Factor: lake_id
    - Factor: station_id
    - Replicate: repeat
    - Response: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lake (lake_id) = 76.4%   station = 11.7%   residual = 11.9%   basin ~ 0%
    Nested: basin > lake > station (REML)

>> COMMENTARY (narration):
    We partitioned the variability of chlorophyll by the spatial scale it comes from with variance-components
    analysis: basin, lake and station levels are nested. The result is striking: 76% of the variance comes from
    between-lake differences, only 12% from between-station (within-lake) differences, and the basin level is nearly
    zero. So "lake identity" is the most determining scale -- lakes differ greatly, but different stations of the same
    lake are similar. This directly guides monitoring design: monitoring many lakes with few stations is more
    informative than few lakes with many stations. VARCOMP optimises the sampling strategy by answering "where is the
    variability".

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_heatwave_summer_chla.xlsx
  >> SCENARIO (narration):
    In this study we ask whether a warm, fluctuating summer regime changes the
    algal biomass of a shallow lake compared to a normal summer. We grouped
    sampling occasions by summer_type into normal and warm_fluctuating
    conditions and measured chlorophyll-a (chla_ugL) as our indicator of
    phytoplankton biomass. A Bayesian t-test fits perfectly here because it lets
    us quantify the evidence for a real difference between the two summer types
    and express our uncertainty directly as a posterior, rather than just
    rejecting or failing to reject a null. This way we can say how credible it
    is that heatwave-like summers elevate chla in the lake.
  >> VARIABLE SELECTION:
    - Group: summer_type
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 4.14e+26 (decisive evidence)   Cohen's d = 2.60
    Heatwave summers vs normal summers -- chlorophyll

>> COMMENTARY (narration):
    We tested whether chlorophyll differs between heatwave summers and normal summers with a Bayesian t-test. The
    Bayes Factor BF10 = 4.14e+26 -- astronomically large, "decisive evidence": the data support the hypothesis of a
    difference quadrillions of times more than no difference. Cohen's d = 2.60 makes the effect enormous. Unlike a
    classic p-value, the Bayes Factor directly measures the strength of evidence for both H1 and H0 and distinguishes
    "absence of evidence" from "evidence of absence". The Bayesian framework convincingly shows how strongly the
    climate-temperature effect reflects on lake productivity.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_chla_temperature_BF10.xlsx
  >> SCENARIO (narration):
    Here we investigate how tightly chlorophyll-a tracks water temperature
    across our lake monitoring samples. We pair each observation's temperature_C
    with its chla_ugL to ask whether warmer water is associated with higher
    algal biomass. A Bayesian correlation is ideal because it returns a Bayes
    factor (BF10) that directly weighs the evidence for an association against
    the evidence for no association, alongside a credible interval for the
    correlation coefficient. This gives a richer picture of the temperature-
    chlorophyll link than a single p-value.
  >> VARIABLE SELECTION:
    - Variable 1: temperature_C
    - Variable 2: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.735 (very strong)   BF10 = 4.75e+08 (decisive evidence)
    Chlorophyll - temperature relationship

>> COMMENTARY (narration):
    We assessed the relationship between chlorophyll and temperature with Bayesian correlation: r = 0.74 is very
    strong, and the Bayes Factor BF10 = 4.75e+08 makes the evidence for the relationship's existence decisive. As
    temperature rises, chlorophyll (algal production) clearly increases -- a direct indicator of warming-accelerated
    eutrophication. Unlike classic correlation, Bayesian correlation expresses the strength of the relationship with
    a probability distribution and an evidence factor -- also answering "how sure are we". It is a powerful evidence
    framework for studies quantifying climate change's effect on lake ecosystems.

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_nutrient_salinity_chla.xlsx
  >> SCENARIO (narration):
    This experiment examines how nutrient enrichment and salinity jointly shape
    phytoplankton biomass in mesocosms. We crossed two nutrient levels (low_N,
    high_N) with three salinity levels (fresh, mild_saline, saline) and recorded
    chlorophyll-a (chla_ugL) in each unit. A Bayesian two-way ANOVA is the right
    tool because it lets us compare models with and without main effects and
    their interaction, quantifying the evidence that nutrients, salinity, or
    their combination drive chla. The posterior also tells us the credible size
    of each effect on algal biomass.
  >> VARIABLE SELECTION:
    - Factor 1: nutrient
    - Factor 2: salinity
    - Outcome: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    nutrient: F(1,118) = 382.54, eta2p = 0.76, BF10 = 4.11e+36 (decisive evidence)
    Nutrient x salinity -> chlorophyll

>> COMMENTARY (narration):
    We examined the effect of nutrient and salinity factors on chlorophyll with Bayesian ANOVA. The nutrient effect
    is overwhelming: F(1,118) = 382.54, effect size eta2p = 0.76 (three-quarters of the variance!), Bayes Factor
    4.11e+36 -- decisive evidence. So phosphorus/nitrogen input is the dominant factor determining chlorophyll.
    Bayesian ANOVA gives a separate Bayes Factor for each effect and interaction, answering "which factor really
    matters" more informatively than classic ANOVA -- and can even provide "evidence of absence" for non-significant
    effects. It is a powerful framework for experimental nutrient-manipulation studies.

====================================================================================

#104  Bayesian Hierarchical
    file: 104_hierarchical_bayesian_country_mesocosm_LMM.xlsx
  >> SCENARIO (narration):
    Across a pan-European mesocosm network we want to know how total phosphorus
    and temperature influence algal biomass while respecting that observations
    are nested within countries and individual mesocosms. We model chlorophyll-a
    (chla_ugL) as a function of TP_ugL and temperature_C, with random intercepts
    for country and mesocosm_id to absorb regional and unit-level variation. A
    Bayesian hierarchical (multilevel) model fits because it partially pools
    information across countries and yields full posterior distributions for
    both the nutrient and temperature slopes. This lets us generalise the
    warming and eutrophication effects while honestly accounting for the grouped
    design.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: TP_ugL
    - Predictor: temperature_C
    - Grouping: country
    - Grouping: mesocosm_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group (country) random variance = 7.30 (20.2%)   ICC = 0.202  [95% CrI: 0.155, 0.249]
    Chlorophyll ~ phosphorus + temperature  +  (1 | country)   (Bayesian)

>> COMMENTARY (narration):
    We analysed the pan-European mesocosm data with a Bayesian hierarchical model -- the Bayesian counterpart of LMM.
    The country-level variance is 20% of the total (ICC = 0.20), reported with a credible interval (CrI). The strength
    of the Bayesian hierarchical model: it expresses both within- and between-group uncertainty with full probability
    distributions and balances groups with few observations via "partial pooling" -- softening extreme estimates. In
    multi-country, multi-centre or nested ecological data, it is the modern way to estimate both fixed and random
    effects convincingly.

====================================================================================

#105  Spatial Lag (SAR)
    file: 105_spatial_sar_konya_lake_TP_chla.xlsx
  >> SCENARIO (narration):
    For a cluster of lakes in the Konya basin we ask how nutrients and catchment
    land use drive chlorophyll-a, while accounting for the fact that nearby
    lakes tend to resemble one another. We model chla_ugL using TP_ugL, TN_mgL,
    temperature_C and agriculture_percent, with lat and lon defining the spatial
    neighbourhood structure. A spatial lag (SAR) model is appropriate because
    the chla of one lake may depend on the chla of its neighbours, and ignoring
    this autocorrelation would bias the nutrient coefficients. The spatial lag
    term captures this neighbour spillover in algal biomass.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: TP_ugL
    - Predictor: TN_mgL
    - Predictor: temperature_C
    - Predictor: agriculture_percent
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho (spatial lag) = 0.065   (weak positive neighbour effect)
    Chlorophyll ~ phosphorus + nitrogen + temperature + agriculture%  +  neighbour chlorophyll

>> COMMENTARY (narration):
    Modelling chlorophyll in the Konya basin lakes, we accounted for spatial dependence -- the tendency of nearby
    lakes to be similar. The SAR (spatial autoregressive) model links a lake's chlorophyll to its neighbours'
    chlorophyll too (the rho parameter). Here rho = 0.065, weak positive: a neighbour effect exists but is small;
    chlorophyll is mainly set by local environmental variables (phosphorus, agriculture). Using a spatial model is
    critical, because ordinary regression gives biased/spuriously significant results when "spatial autocorrelation"
    is present. SAR models spatial processes where "neighbour values also have an effect" (diffusion, spillover).

====================================================================================

#106  Spatial Error
    file: 106_spatial_error_konya_lake_lambda.xlsx
  >> SCENARIO (narration):
    Using the same Konya lake network, we again model chlorophyll-a from
    nutrients, temperature and surrounding agriculture, but now we suspect the
    spatial dependence lies in unmeasured shared drivers rather than in
    neighbours' biomass. We regress chla_ugL on TP_ugL, TN_mgL, temperature_C
    and agriculture_percent, with lat and lon defining spatial proximity. A
    spatial error model fits because it places the autocorrelation in the
    residuals via a lambda term, correcting our standard errors for clustered
    omitted factors like shared climate or hydrology. This gives unbiased
    inference on how nutrients and land use relate to algal biomass.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: TP_ugL
    - Predictor: TN_mgL
    - Predictor: temperature_C
    - Predictor: agriculture_percent
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda (spatial error) = -0.223
    Chlorophyll ~ environmental variables  +  spatial error structure

>> COMMENTARY (narration):
    The spatial error model addresses a different kind of spatial dependence than SAR: not a DIRECT effect of
    neighbour values, but correlation induced through the residuals by unobserved (omitted) spatial factors. The
    lambda parameter (-0.223) measures this spatial error structure. Which spatial model is appropriate (SAR or SEM)
    depends on whether the effect comes "from neighbour values or from unmeasured common spatial factors". Either
    way, ignoring spatial structure leads to error. SEM cleans out the effect of unmeasured spatial confounders
    (geology, climate gradient) to estimate the true effect of environmental variables more accurately.

====================================================================================

#107  GWR
    file: 107_gwr_konya_local_TP_chla_relationship.xlsx
  >> SCENARIO (narration):
    Here we explore whether the strength of the phosphorus-chlorophyll
    relationship varies from place to place across our lakes rather than being
    constant. We model chla_ugL against TP_ugL, TN_mgL, temperature_C and
    agriculture_percent, but allow the coefficients to change with location
    using lat and lon. Geographically Weighted Regression (GWR) is the natural
    choice because it fits a local regression around each lake, mapping where
    nutrients most strongly control algal biomass. This reveals spatial non-
    stationarity that a single global model would hide.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Predictor: TP_ugL
    - Predictor: TN_mgL
    - Predictor: temperature_C
    - Predictor: agriculture_percent
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.891   (local/location-specific regression)
    Spatial variation of the phosphorus-chlorophyll relationship

>> COMMENTARY (narration):
    Ordinary regression assumes a SINGLE phosphorus-chlorophyll relationship for the whole region; but this
    relationship can vary in space -- phosphorus may be very influential in one sub-basin and weak in another. GWR
    (Geographically Weighted Regression) captures this spatial heterogeneity by fitting a separate local regression
    at each location (R2 = 0.89). The output is a map of coefficients: a surface showing where phosphorus's effect on
    chlorophyll is strong and where it is weak. This tests "is the relationship the same everywhere" and sets local
    management priorities -- where phosphorus control helps most. In spatially varying processes, GWR reveals what
    the global model hides.

====================================================================================

#108  Nested LMM
    file: 108_nested_lmm_lake_station_date_chla.xlsx
  >> SCENARIO (narration):
    This mesocosm experiment has a strictly hierarchical layout: mesocosm units
    sit within countries, which sit within experimental blocks. We analyse
    chlorophyll-a (chla_ugL) while accounting for this nesting of mesocosm_no
    within country within block. A nested linear mixed model is required because
    the random intercepts must reflect the containment structure of the design,
    partitioning variance in algal biomass across blocks, countries and
    individual mesocosms. This prevents pseudoreplication and isolates the true
    residual variability.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Nesting level: block
    - Nesting level: country
    - Nesting level: mesocosm_no

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    country (P): F = 55.66, p < .001 ***   |   mesocosm x country (RxP): F = 7.66, p < .001 ***
    block [within country] (F): F = 3.82, p < .001 ***

>> COMMENTARY (narration):
    In this design the measurements are nested: country > block > mesocosm. The nested mixed model tests each level's
    contribution to the variance separately. The results are significant at all levels: between-country differences
    are largest (F = 55.66), but there are real differences at the block and mesocosm levels too. Mixing up the
    levels in nested data (e.g. ignoring country) produces spurious significance or wrong standard errors. Nested LMM
    is the correct framework for hierarchical sampling designs (country/site/replicate, school/class/student),
    separating the genuine contribution of each scale.

====================================================================================

#109  Crossed LMM
    file: 109_crossed_lmm_climate_nutrient_mesocosm.xlsx
  >> SCENARIO (narration):
    In this climate-by-nutrient mesocosm trial, the same mesocosm identities
    recur across different climate regions and nutrient treatments, so the
    grouping factors are crossed rather than nested. We model chlorophyll-a
    (chla_ugL) with climate_region and nutrient_loading as effects of interest,
    while treating mesocosm_id as a crossed random factor that cuts across both.
    A crossed linear mixed model is the right framework because it
    simultaneously accounts for variability tied to climate region and to the
    shared mesocosm units. This yields clean estimates of how warming and
    nutrient loading affect algal biomass.
  >> VARIABLE SELECTION:
    - Outcome: chla_ugL
    - Fixed factor: climate_region
    - Fixed factor: nutrient_loading
    - Crossed random: mesocosm_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    climate region: F = 1886.96, p < .001 ***   |   nutrient loading: F = 3458.97, p < .001 ***
    climate x nutrient interaction: F = 0.16, p = 0.856 (non-significant)

>> COMMENTARY (narration):
    Unlike nesting, here two factors are crossed: each climate region is combined with each nutrient level (fully
    factorial). The crossed mixed model handles this. Both the climate (F = 1887) and nutrient (F = 3459) main
    effects are overwhelmingly significant -- both strongly affect chlorophyll. The interaction is non-significant
    (p = 0.86): the nutrient effect is similar across all climate regions, and the climate effect similar across all
    nutrient levels -- the effects are additive, not amplifying each other. This answers "is there an interaction
    when two factors are tested together"; it is the correct analysis for crossed experimental designs (climate x
    nutrient mesocosm experiments).

====================================================================================

#110  KDE Density Map
    file: 110_KDE_lake_sampling_density.xlsx
  >> SCENARIO (narration):
    We want to visualise where our lake sampling effort is concentrated across
    the study region and how it overlaps with algal hotspots. Using the lat and
    lon of each sampling unit, we build a kernel density surface of sampling
    locations, and we can weight or interpret it alongside chla_ugL to see
    whether dense sampling coincides with high chlorophyll areas. A KDE density
    map is ideal because it smooths discrete points into a continuous intensity
    surface, highlighting clusters of monitoring activity. This helps assess
    spatial coverage and detect under-sampled parts of the lake network.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Weight: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    95 lake points   mean chlorophyll = 36.0   latitude range 37.1 - 40.9
    Sampling-density map via kernel density estimation

>> COMMENTARY (narration):
    On the Map tab, we render the geographic density of lake sampling points into a heat map with KDE (Kernel Density
    Estimation). KDE turns points into a smooth density surface: we can see where sampling/lake clustering is dense
    and where there are gaps. The distribution of 95 lakes shows how monitoring effort concentrates geographically.
    This is used both to detect sampling bias (which regions are under-represented) and to see high-lake-density "hot
    regions". It is the basic mapping tool for visualising spatial point data.

====================================================================================

#111  Hexbin Map
    file: 111_Hexbin_pondscape_participant.xlsx
  >> SCENARIO (narration):
    We surveyed 150 sampling units across a pondscape and recorded the
    geographic position of each one with lat and lon, while a field crew
    classified the local sampling density into a categorical density_interval of
    low, medium, or high. To communicate where sampling effort and pond
    clustering concentrate across the landscape, we want to bin the many
    overlapping points into clear spatial cells rather than plotting them
    individually. A Hexbin Map aggregates the lat and lon coordinates into
    hexagonal cells, revealing the spatial structure of the density_interval
    pattern at a glance. This is the right choice when point overplotting hides
    the true geographic distribution of our limnological sampling units.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Category: density_interval

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    150 participant/pond points   aggregated into hexagonal cells
    Each cell shows the point density within it by colour

>> COMMENTARY (narration):
    The hexbin map aggregates many points into hexagon cells and shows each cell's density by colour -- a clear
    solution, alternative to KDE, when thousands of points overlap into an unreadable mess. We rendered 150 pond
    participant points into a hexagonal grid to visualise where participation/observation is dense. Hexagonal cells
    carry less edge-bias than squares and represent neighbour relations more evenly. It is the practical way to turn
    dense point data (survey participation, species observations, event locations) into a readable density map.

====================================================================================

#112  Moran's I
    file: 112_Morans_I_lake_chla_autocorrelation.xlsx
  >> SCENARIO (narration):
    At 80 lake monitoring stations we measured chlorophyll-a concentration
    (chla_ugL) together with each station's geographic coordinates lat and lon.
    We want to know whether eutrophication is spatially structured: do stations
    with high chla_ugL sit next to other high stations, or is the algal biomass
    scattered randomly across the lake? Moran's I tests for global spatial
    autocorrelation, quantifying whether chla_ugL values at nearby locations are
    more similar than expected by chance. This fits because it directly answers
    whether chlorophyll-a clusters geographically rather than being randomly
    distributed.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Value: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.531   E[I] = -0.013   z = 9.28   p < .001
    Spatial autocorrelation of chlorophyll (strong positive)

>> COMMENTARY (narration):
    Moran's I is the global spatial-autocorrelation statistic measuring "do nearby lakes have similar chlorophyll".
    The result is strongly positive (I = 0.53 against an expected -0.01, z = 9.28, p < .001): chlorophyll is
    spatially clustered -- high-chlorophyll lakes lie near each other, and so do low ones. This is not chance; shared
    basin, climate and land use make neighbouring lakes similar. The methodological importance is large: if spatial
    autocorrelation exists, ordinary statistics are wrong (hence the spatial models in #105-107 are needed). Moran's
    I is the starting diagnostic of spatial analysis.

====================================================================================

#113  Getis-Ord Hotspot
    file: 113_Getis_Ord_chla_eutrophic_hotspot.xlsx
  >> SCENARIO (narration):
    Across 68 lake sampling points we recorded chlorophyll-a concentration
    (chla_ugL) and the geographic location of each point with lat and lon.
    Knowing that elevated chla_ugL signals eutrophic conditions, we want to
    pinpoint exactly where statistically significant clusters of high algal
    biomass form so managers can target those zones. The Getis-Ord Gi* hotspot
    analysis evaluates each location against its neighbors to flag significant
    hot and cold spots of chla_ugL. This is the appropriate test because it
    localizes eutrophic hotspots rather than just confirming that clustering
    exists overall.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Value: chla_ugL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    28 hot spots (high-chlorophyll cluster)   28 cold spots (low-chlorophyll cluster)   n = 68
    Getis-Ord Gi* local statistic (z > 1.96 significant)

>> COMMENTARY (narration):
    While Moran's I reports the overall clustering of the whole region, Getis-Ord Gi* shows WHERE the hot/cold spots
    are, location by location. 28 lakes are statistically significant "hot spots" (high chlorophyll surrounded by
    high chlorophyll -- eutrophic clusters), and 28 lakes are "cold spots" (clean-water clusters). This gives direct
    guidance for pollution management: it is most efficient to focus intervention on the hot-spot clusters. Gi*
    hot-spot analysis is the standard spatial method for finding "where the problem concentrates geographically" --
    from crime to disease, from eutrophication to forest fire.

====================================================================================

#114  DBSCAN Spatial
    file: 114_DBSCAN_lake_cluster_coast.xlsx
  >> SCENARIO (narration):
    We collected 115 lake and coastal sampling units, each georeferenced by lat
    and lon, and we want to discover natural groupings of stations purely from
    their spatial arrangement. Rather than fixing the number of groups in
    advance, we need a method that finds dense clusters and isolates outlying
    stations that do not belong to any cluster. DBSCAN groups the lat and lon
    coordinates by spatial density, automatically identifying clusters of
    monitoring sites and labeling sparse points as noise. This density-based
    approach fits because the natural spatial clustering of our sampling units
    is unknown and may include irregular shapes and outliers.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   noise (isolated) = 12   n = 115
    Lake locations clustered by density

>> COMMENTARY (narration):
    On the map we clustered lakes' geographic locations by density with DBSCAN: three dense lake groups were found,
    and 12 lakes were flagged as "isolated/noise" points belonging to no cluster. DBSCAN's strength is that it does
    not require the number of clusters in advance and can capture irregularly shaped clusters -- and automatically
    separates sparse/distant lakes. The three clusters show that lakes group in particular basins/regions. This
    spatial clustering is a practical mapping tool for dividing monitoring networks into regions, defining management
    units, and assessing geographically similar lakes together. With this we complete the full 114-analysis tour of
    the Limnology package.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_chlorophyll_a_ugL.xlsx
  >> SCENARIO (narration):
    We follow 40 stations measured at three season levels (spring, summer, autumn); each belongs to one of two trophic_state groups (eutrophic / oligotrophic). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on chlorophyll-a.
  >> VARIABLE SELECTION:
    - Dependent variable: chlorophyll_a_ugL
    - Subject ID: station_id
    - Between-subjects factor: trophic_state
    - Within-subjects factor: season

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (trophic_state): F(1,38) = 126.64  p < .001  np2 = 0.769
    Within (season):  F(2,76) = 36.63  p < .001  np2 = 0.491
    Interaction:         F(2,76) = 15.60  p < .001  np2 = 0.291
    Mauchly W = 0.897  p = 0.128   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a trophic_state x season mixed design we analyzed chlorophyll-a for 40 stations (120 observations). The interaction is significant (F(2,76) = 15.60, p < .001, np2 = 0.291) *** -- the two groups' change across season differs in magnitude. The between-subjects main effect (eutrophic vs oligotrophic) is F = 126.64, p < .001; the within-subjects main effect (spring/summer/autumn) is F = 36.63, p < .001. Mauchly's test p = 0.128, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each season level. In limnology, the mixed design is the standard analysis for seasonal monitoring of lakes of differing trophic state.

====================================================================================
