====================================================================================
MerQur - Zeki_Kaya_Orman_Genetigi - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Zeki_Kaya_Orman_Genetigi/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 001_descriptive_pinus_brutia_provenance_trial.xlsx
  >> SCENARIO (narration):
    In this provenance trial we established 280 Pinus brutia seedlings drawn
    from six native seed sources, ranging from Antalya_Duzlercami and
    Mut_Karaman to Adana_Pos, and we want a clear first picture of how this
    material performs before any formal comparison. We summarize the early
    growth and adaptive traits, including height_cm, root_neck_diameter_mm,
    height_growth_annual_cm and side_branch_count, alongside phenology and
    stress indicators such as bud_fracture_day, cold_damage_score_1_5 and
    drought_damage_score_1_5. We also describe wood-related traits like
    tree_specific_density_gcm3 and dry_item_percent across the different
    provenance groups. Descriptive statistics are the right starting point here
    because they reveal central tendency, spread and the overall distribution of
    each trait, guiding which formal tests come next.
  >> VARIABLE SELECTION:
    - Group: provenance
    - Measure: seedling_age_year
    - Measure: height_cm
    - Measure: root_neck_diameter_mm
    - Measure: height_growth_annual_cm
    - Measure: side_branch_count
    - Measure: bud_fracture_day
    - Measure: cold_damage_score_1_5
    - Measure: drought_damage_score_1_5
    - Measure: tree_specific_density_gcm3
    - Measure: dry_item_percent

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 trees (Pinus brutia provenance trial)
    height : mean = 120.4 cm   |   root-neck diameter : mean = 17.6 mm   |   wood specific density : mean = 0.485 g/cm3

>> COMMENTARY (narration):
    We first drew the general profile of the provenance (origin) trial: 280 Turkish red pine seedlings, mean height
    120 cm, root-neck diameter 17.6 mm, wood specific density 0.485 g/cm3. This descriptive table sets the stage for
    all the genetic analyses that follow -- differences among provenances, heritability, G×E interaction. In forest
    genetics, before any inferential test, summarising the trial material's basic growth and quality traits is
    essential both to check data quality and to define breeding goals.

====================================================================================

#2  Normality Tests
    file: 002_normality_wood_specific_gravity_PinusNigra.xlsx
  >> SCENARIO (narration):
    For 220 Pinus nigra trees from various half_sib_family lines, we measured
    wood specific gravity (wsg_gcm3) together with dbh_cm and tree_age_year, and
    before fitting parametric models we must confirm the distributional
    assumptions. The key question is whether wsg_gcm3 follows a normal
    distribution, since many downstream genetic and growth analyses rely on this
    assumption. We run normality tests on the wood quality and growth variables
    to decide between parametric and non-parametric approaches. This step is
    essential because violating normality can bias heritability estimates and
    trait comparisons.
  >> VARIABLE SELECTION:
    - Test: wsg_gcm3
    - Test: dbh_cm
    - Test: tree_age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pinus nigra, 3 continuous variables tested:
    wsg_gcm3 (wood density) : Shapiro-Wilk = 0.998  p = 0.993   Normal
    dbh_cm (breast-height diameter) : n = 200   mean = 29.06   skew = 0.16   kurt = -0.08
                                      Shapiro-Wilk = 0.9954  p = 0.804   KS = 0.0343  p = 0.826   D'Agostino p = 0.643   NORMAL
    tree_age_year (age)     : Shapiro-Wilk = 0.948  p < .001   NOT NORMAL

>> COMMENTARY (narration):
    We tested whether three continuous variables in black pine half-sib families -- wood specific gravity, breast-
    height diameter (dbh) and tree age -- are normally distributed. The result splits instructively: wood density and
    dbh are nearly perfectly normal (all three tests p > 0.05; dbh skew 0.16, kurtosis -0.08, both near zero), but
    tree age deviates significantly from normality (p < .001). This is an expected pattern -- growth/quality traits
    (diameter, density) are the sum of many small genetic+environmental effects, hence near-normal (central limit
    theorem); age may be skewed/multi-modal depending on the trial design (discrete age classes). The practical
    takeaway: we can safely use parametric tests (t-test, ANOVA, Pearson, heritability estimation) on dbh and wsg; for
    age, nonparametric methods or a transform are more appropriate. The normality check decides, for each variable
    separately, which test family is appropriate.

====================================================================================

#3  One-Sample t-Test
    file: 003_one_sample_t_Liquidambar_He_vs_reference.xlsx
  >> SCENARIO (narration):
    Across 30 Liquidambar populations we estimated the mean expected
    heterozygosity (mean_heterozygosity_He) and want to know whether genetic
    diversity in our sampled stands meets an established reference value for the
    species. We compare the observed mean_heterozygosity_He against the fixed
    reference stored in test_obtained_reference_He to judge whether our
    populations are more or less diverse than expected. A one-sample t-test fits
    perfectly because we are testing a single continuous variable against a
    known constant. The outcome tells conservation managers whether these
    populations harbor adequate genetic variation.
  >> VARIABLE SELECTION:
    - Test variable: mean_heterozygosity_He

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(29) = -4.9077   p < .001 ***   Cohen d = -0.896 (Large)
    Mean He = 0.334   (test mu = 0.40, reference heterozygosity)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the mean heterozygosity (He, a genetic-diversity measure) of Liquidambar (sweetgum) populations
    against a reference value of 0.40. The result is concerning: mean He is 0.334, significantly below the reference
    -- t(29) = -4.91, p < .001, large effect size d = -0.90. So these populations' genetic diversity is below the
    reference by more than chance can explain. This matters for conservation genetics: low He points to a population
    that may have undergone a bottleneck or shrunk, with limited future adaptive potential. The one-sample t-test is
    the right way to compare a genetic indicator against a known reference.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent Samples t-Test
    file: 004_independent_t_PinusBrutia_Toros_mediterranean_height.xlsx
  >> SCENARIO (narration):
    We compare 120 Pinus brutia trees split into two provenance_group origins,
    Toros versus mediterranean, to see whether seed source affects early growth.
    The response is height_cm_3yr, three-year height, which we expect to differ
    between these contrasting climatic origins. An independent samples t-test is
    appropriate because we have one continuous outcome and two unrelated groups
    of trees. The result will indicate which provenance establishes faster and
    is better suited for reforestation.
  >> VARIABLE SELECTION:
    - Group: provenance_group
    - Test variable: height_cm_3yr

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(118) = -4.030   p < .001 ***   Cohen d = -0.736 (Moderate)
    Height differs significantly by provenance group.   H0 REJECTED

>> COMMENTARY (narration):
    We compared red pine seedling height between two provenance groups (Taurus vs Mediterranean) with an independent
    t-test. The difference is significant and moderate-large (t(118) = -4.03, p < .001, d = -0.74): one provenance
    group's seedlings are markedly taller. This is direct evidence of genetic differentiation among provenances --
    where the tree comes from (seed source) determines growth performance. Provenance selection is the most critical
    decision in afforestation; this result is its statistical basis. The independent t-test is the basic method for
    comparing two provenance/group means.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired t-Test
    file: 005_paired_t_saffron_corm_dokukulturu_8wk.xlsx
  >> SCENARIO (narration):
    In a tissue-culture experiment with 36 jars, we counted corms before
    treatment (control_before_corm_count) and again after eight weeks
    (8wk_post_corm_count) to evaluate corm multiplication. Because the two
    measurements come from the same jar, they are naturally paired, and we want
    to know whether corm count increased significantly over the eight-week
    period. A paired t-test is the correct choice since each before value is
    linked to its own after value. This tells us whether the propagation
    protocol effectively boosts corm production.
  >> VARIABLE SELECTION:
    - Before: control_before_corm_count
    - After: 8wk_post_corm_count

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Paired t(35) = -20.2701   p < .001 ***   Cohen d_z = -3.378 (Large)
    Mean difference (before − 8-week post corm count) = -3.47   H0 REJECTED

>> COMMENTARY (narration):
    In saffron (Crocus) tissue culture, we compared the initial and 8-week corm (tuber) count in the same jars with a
    paired t-test. The result is striking: corm count rose by an average of 3.47 over 8 weeks, t(35) = -20.27,
    p < .001, with an enormous effect size (d_z = -3.38). The tissue-culture protocol drove very strong multiplication.
    Because we measured the same jars twice, the paired test is correct -- it removes between-jar variability and
    focuses only on the within-jar change. It is the right way to evaluate "before/after" in plant tissue-culture/
    micropropagation studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 006_one_way_anova_PinusBrutia_5provenance_height.xlsx
  >> SCENARIO (narration):
    We grew 150 Pinus brutia seedlings representing five provenances, Antalya,
    Mugla, Mersin, Burdur and Kahramanmaras, and recorded three-year height
    (height_cm_3yr) to test for provenance differences in growth. The research
    question is whether mean height varies significantly among these five seed
    sources. A one-way ANOVA is ideal because we compare one continuous response
    across more than two groups simultaneously. Identifying the top-performing
    provenance supports selection of superior planting stock.
  >> VARIABLE SELECTION:
    - Factor: provenance
    - Test variable: height_cm_3yr

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(4, 145) = 12.9188   p < .001 ***   η² = 0.2627
    DECISION: H0 REJECTED (5 provenances)

>> COMMENTARY (narration):
    We compared the 3-year height of five red pine provenances with one-way ANOVA. The result is significant
    (F(4,145) = 12.92, p < .001), with effect size η² = 0.26 -- more than a quarter of height variance is explained by
    provenance, a strong effect in forest genetics. ANOVA says at least one provenance differs from the others; a
    post-hoc test is needed to see which. "Which provenance grows better" is the basic question of provenance trials
    and is decisive for seed transfer zones/breeding selection; ANOVA is the standard way to answer it across multiple
    groups.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 007_two_way_anova_provenance_site_GxE_height.xlsx
  >> SCENARIO (narration):
    To study genotype-by-environment interaction, we planted four provenances
    (Antalya, Mugla, Mersin, Burdur) across three trial sites (Ankara,
    Eskisehir, Konya) and measured height_cm on 96 trees. We ask not only
    whether provenance and site each affect height_cm, but also whether
    provenance ranking changes across sites, which is the essence of GxE. A two-
    way ANOVA is the right design because it tests two crossed factors and their
    interaction on a single continuous outcome. Detecting interaction guides
    whether we recommend provenances regionally or broadly.
  >> VARIABLE SELECTION:
    - Factor 1: provenance
    - Factor 2: site
    - Test variable: height_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    provenance (main) : F(3) = 12.17   p < .001 ***   η²p = 0.30
    site (main)       : F(2) =  5.17   p = 0.008 **    η²p = 0.11

>> COMMENTARY (narration):
    We examined height with two factors -- provenance (genotype) and trial site (environment) -- together; this is the
    Genotype×Environment (G×E) question at the heart of forest genetics. The provenance effect is strong (F = 12.17,
    η²p = 0.30), the site effect also significant (F = 5.17, p = 0.008, η²p = 0.11). So both genetic origin and growing
    environment determine height. Two-way ANOVA shows both factors' effects -- separately and jointly (interaction).
    If the interaction is significant, the question "is the best provenance the same at every site, or site-dependent"
    is answered -- a core inference for seed-transfer policy. G×E is critical for provenance selection under climate change.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated Measures ANOVA
    file: 008_RM_anova_4yil_height_PinusBrutia.xlsx
  >> SCENARIO (narration):
    We followed 40 Pinus brutia seedlings from four provenances, including
    Mugla_Marmaris and Mersin_Erdemli, recording height each year across
    year1_height_cm through year4_height_cm. The question is how height develops
    over four consecutive years and whether the growth trajectory differs among
    provenances. A repeated measures ANOVA is appropriate because the same
    seedlings are measured repeatedly over time, creating correlated
    observations. This reveals both the overall growth trend and provenance-
    specific differences in growth rate.
  >> VARIABLE SELECTION:
    - Within-subject: year1_height_cm
    - Within-subject: year2_height_cm
    - Within-subject: year3_height_cm
    - Within-subject: year4_height_cm
    - Between-subject: provenance

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 117) = 506.6967   p < .001 ***   η²p = 0.9285   (n = 40)
    DECISION: H0 REJECTED

>> COMMENTARY (narration):
    We compared the same 40 seedlings' height across four consecutive years with repeated-measures ANOVA. Because the
    same seedlings are measured repeatedly, observations are dependent; RM-ANOVA accounts for this within-subject
    correlation. The result is overwhelming (F(3,117) = 506.70, p < .001, η²p = 0.93): seedlings grow very markedly
    over the years -- an expected but quantified fact (effect size 93%!). RM-ANOVA is the correct design for
    longitudinal growth trials where the same trees are tracked over time, and is far more powerful than treating each
    year as independent. It is ideal for studies following tree growth curves.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 009_manova_provenance_4kantitatif_feature.xlsx
  >> SCENARIO (narration):
    For 88 trees from four provenances (Antalya, Mugla, Mersin, Burdur) we
    measured four quantitative traits at once: height_cm, diameter_mm,
    branch_count and wsg_gcm3. Rather than testing each trait separately, we ask
    whether the provenances differ when these correlated traits are considered
    jointly as a multivariate profile. A MANOVA fits because it tests group
    differences across several dependent variables simultaneously while
    accounting for their correlations. This gives a holistic view of how
    provenances differ in their overall growth-and-wood phenotype.
  >> VARIABLE SELECTION:
    - Factor: provenance
    - Dependent: height_cm
    - Dependent: diameter_mm
    - Dependent: branch_count
    - Dependent: wsg_gcm3

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.572 -> F(12, 214) = 4.20   p < .001 ***   (n = 88)
    Dependent: height, diameter, branch count, wood density   |   Factor: provenance

>> COMMENTARY (narration):
    Here we have four correlated quantitative traits -- height, diameter, branch count, wood density -- and one
    provenance factor. Running four separate ANOVAs would inflate the error rate and ignore the correlations among
    traits. MANOVA tests all four jointly: Wilks' Lambda F = 4.20, p < .001, so provenance strongly shifts the
    multivariate phenotype profile. MANOVA shows that provenances differ jointly across several related traits -- in
    tree breeding, selection is usually based not on a single trait but on a combination. After a significant MANOVA,
    follow-up univariate tests show which trait drives the difference.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 010_ancova_clone_wsg_dbh_covariance.xlsx
  >> SCENARIO (narration):
    Across 150 trees from several clones we want to compare wood specific
    gravity (wsg_gcm3) among clones, but stem size measured as dbh_cm is known
    to influence wood density and could confound the comparison. We therefore
    test for clone differences in wsg_gcm3 while statistically controlling for
    dbh_cm as a covariate. An ANCOVA is the appropriate method because it
    adjusts the group means for a continuous covariate before comparing them.
    This yields a fairer clonal comparison of intrinsic wood quality independent
    of tree size.
  >> VARIABLE SELECTION:
    - Factor: clone
    - Dependent: wsg_gcm3
    - Covariate: dbh_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    clone (factor): η²p = 0.325 (large)   p < .001 ***
    Covariate: dbh_cm (adjusted for stem diameter)   |   DV: wsg_gcm3   H0 REJECTED

>> COMMENTARY (narration):
    When comparing clones' wood density, it is essential to control for stem diameter (dbh) -- because wood density
    can differ in thicker trees, a confounder. ANCOVA takes dbh as a covariate and holds it statistically constant.
    The result: even after adjusting for diameter, clones differ significantly in wood density (η²p = 0.33, large
    effect). So the observed difference is not a by-product of diameter differences; it is the genuine genetic effect
    of clone. In wood-quality breeding, comparing clones at constant diameter is critical -- ANCOVA does this fairly.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 011_bootstrap_CI_Fst_small_N_population.xlsx
  >> SCENARIO (narration):
    When estimating genetic differentiation among small forest populations, the
    number of populations and loci is often limited, so a single Fst point
    estimate can be unreliable. In this study we ask how stable the Fst_paragon
    values are across our sampled populations, given that each population was
    characterized by a different sample_count and a different SSR_locus_count.
    By resampling the data with a Bootstrap Confidence Interval, we generate an
    empirical distribution of Fst and obtain robust interval bounds without
    assuming normality. This is exactly the situation a bootstrap is built for:
    small sample size, no distributional guarantees, and the need to quantify
    uncertainty around a differentiation index.
  >> VARIABLE SELECTION:
    - Group: population
    - Statistic: Fst_paragon
    - Covariate: sample_count
    - Covariate: SSR_locus_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean Fst = 0.195
    95% Bootstrap CI (via resampling)   |   small sample

>> COMMENTARY (narration):
    We derived a confidence interval for the mean of Fst -- which measures genetic differentiation among populations
    -- by bootstrapping: the observed mean Fst is 0.195 (moderate-high differentiation). With few populations/loci,
    classic formulas are unreliable; bootstrap estimates the interval without distributional assumptions, through
    thousands of resamples. It is a robust way to express the uncertainty of population-genetics parameters like Fst.
    An Fst around 0.20 indicates marked genetic structure among populations (likely limited gene flow) -- important
    for conservation and seed-transfer decisions.

====================================================================================

#12  Permutation Test
    file: 012_permutasyon_AMOVA_Salix_populasyonlar.xlsx
  >> SCENARIO (narration):
    We want to test whether Salix populations differ in their molecular genetic
    structure as captured by an AMOVA-style framework. For each subject we
    recorded the two alleles at locus A, allele1_locus_A and allele2_locus_A,
    across a marginal and a center population. Because allele data violate
    normality assumptions and the AMOVA significance is built from random
    shuffling of individuals among groups, we use a Permutation Test to build
    the null distribution by repeatedly reassigning population labels. This non-
    parametric, resampling approach correctly assesses whether the observed
    among-population genetic variance exceeds what we would expect by chance.
  >> VARIABLE SELECTION:
    - Group: population
    - Dependent: allele1_locus_A
    - Dependent: allele2_locus_A

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = 3.14   p = 0.0021 **   H0 REJECTED
    There is a significant difference in allele distribution among populations

>> COMMENTARY (narration):
    We tested allele/genetic differentiation among Salix (willow) populations with a permutation test (AMOVA logic).
    Permutation shuffles the group labels thousands of times to measure whether the observed difference could arise by
    chance -- requiring no distributional assumption, which is why it is standard for molecular-genetic data (AMOVA).
    The result is significant (p = 0.002): populations are genetically more different than chance allows. This
    confirms real genetic structure among populations. AMOVA/permutation is the basic method in population genetics
    for testing how genetic variance is partitioned (among vs within groups).

====================================================================================

#13  Multiple Comparison Corrections
    file: 013_tukey_6Quercus_Hd_pairwise.xlsx
  >> SCENARIO (narration):
    In this work we compare haplotype diversity across six oak species to see
    which taxa harbor the richest chloroplast lineage diversity. The
    haplotype_diversity_Hd was measured for each individual across the type
    groups Q_cerris, Q_robur, Q_ilex, Q_libani, Q_pubescens and Q_petraea.
    Because we will run every pairwise species comparison, the family-wise error
    rate inflates quickly, so we apply Multiple Comparison Corrections such as
    Tukey adjustment to control false positives. This lets us state confidently
    which oak pairs genuinely differ in Hd rather than reporting spurious
    significant differences.
  >> VARIABLE SELECTION:
    - Group: type
    - Dependent: haplotype_diversity_Hd

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Raw (uncorrected): 4/5 significant   |   After correction (Bonferroni/Holm/FDR): 3/5 significant

>> COMMENTARY (narration):
    When we compare the haplotype diversity (Hd) of six Quercus (oak) species pairwise, we run many tests -- each
    carrying a false-positive risk. Multiple-comparison correction controls that risk. Here before correction 4 of 5
    comparisons looked significant, but after correction 3 remain. So 1 "significant" result was actually a false
    positive. This shows why correction is critical in phylogeographic/taxonomic comparisons; skipping it leads to
    reporting spurious diversity differences among species.

====================================================================================

#14  Mann-Whitney U Test
    file: 014_mann_whitney_marginal_center_allele_richness.xlsx
  >> SCENARIO (narration):
    Forest geneticists often expect populations at the range margin to be
    genetically depauperate compared with central populations. Here we test
    whether allele_richness_Ar differs between marginal and center populations
    defined by population_type. Since allelic richness values are bounded and
    frequently skewed rather than normally distributed, a parametric t-test is
    inappropriate for these two independent groups. The Mann-Whitney U Test
    compares the rank distributions and tells us whether marginal stands carry
    significantly lower allele richness.
  >> VARIABLE SELECTION:
    - Group: population_type
    - Dependent: allele_richness_Ar

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Mann-Whitney U = 453.00   p = 0.062 ns   r = .26
    Allele richness does NOT differ significantly by marginal/center population (borderline).   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether allele richness (Ar, a genetic-diversity measure) differs between marginal (range-edge) and
    central populations with the nonparametric Mann-Whitney. The result is borderline non-significant (U = 453,
    p = 0.062, r = 0.26): a difference seems present (moderate effect) but does not cross the 0.05 threshold. The
    "central-marginal hypothesis" -- that central populations are more diverse -- is weakly supported here; a larger
    sample could clarify it. Borderline results should be interpreted cautiously. For ordinal/small-sample genetic
    data Mann-Whitney is more reliable than the t-test.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank Test
    file: 015_wilcoxon_bud_fracture_climate_5yr.xlsx
  >> SCENARIO (narration):
    To track the climatic shift in spring phenology, we recorded bud fracture
    timing for the same seedlings in two years, bud_fracture_2020_jd and
    bud_fracture_2024_jd, expressed as Julian days. Because the measurements are
    paired within each seedling and the day-of-year differences are not
    guaranteed to be normally distributed, a paired t-test is risky. The
    Wilcoxon Signed-Rank Test evaluates whether the median within-seedling
    change in bud fracture date is significant. This reveals whether budburst
    has advanced over the five-year interval under warming conditions.
  >> VARIABLE SELECTION:
    - Measurement 1: bud_fracture_2020_jd
    - Measurement 2: bud_fracture_2024_jd

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilcoxon W = 0.00   p < .001 ***   r = 0.876   n (nonzero) = 26
    Bud-flush day changed significantly between 2020 and 2024.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the bud-flush day (bud fracture, a phenology indicator) in the same seedlings in 2020 and 2024 with
    the paired Wilcoxon test. The result is very strong: W = 0, p < .001, effect size r = 0.88. Bud-flush timing
    changed in the same direction and by a large amount in nearly all seedlings -- direct evidence of climate change's
    effect on tree phenology. For paired measures and ordinal/skewed phenology data, Wilcoxon is the right choice.
    Phenological shifts are among the most sensitive early indicators of climate change's effect on forest ecosystems.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 016_kruskal_wallis_Astragalus_4seksiyon_leaf.xlsx
  >> SCENARIO (narration):
    We investigate whether leaf morphology differs among four taxonomic sections
    of the genus Astragalus. The leaf_length_mm was measured on individuals
    belonging to sect_Hypoglottidei, sect_Stipulati, sect_Onobrychoidei and
    sect_Acanthophace. Leaf length distributions across these sections are
    uneven and not reliably normal, so a one-way ANOVA is not ideal for these
    independent groups. The Kruskal-Wallis Test compares the rank distributions
    of leaf length across the four sections to detect overall morphological
    differentiation.
  >> VARIABLE SELECTION:
    - Group: section
    - Dependent: leaf_length_mm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(3) = 43.92   p < .001 ***   eta-squared_H = 0.38
    Leaf length differs significantly among Astragalus sections.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether leaf length varies among four Astragalus (milkvetch) taxonomic sections with Kruskal-Wallis --
    the nonparametric ANOVA for more than two groups. The result is highly significant (H(3) = 43.92, p < .001), with
    a large effect (eta-squared = 0.38): leaf morphology differs markedly among sections. This is the basis of
    morphological taxonomy -- sections separate by measurable morphological traits. Morphometric data are often not
    normally distributed, so Kruskal-Wallis is more appropriate than classic ANOVA. It is the basic test that
    quantifies species/section distinctions in plant systematics.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 017_friedman_clone_4site_growth_same_clone.xlsx
  >> SCENARIO (narration):
    In a clonal field trial, identical clones were planted at four sites to
    evaluate genotype-by-environment growth responses. For each clone_id we have
    growth recorded at site_Antalya, site_Ankara, site_Eskisehir and site_Konya,
    so the four measurements are repeated on the same genetic individual.
    Because the growth data are not normally distributed and the design is a
    repeated-measures block, the Friedman Test is the appropriate non-parametric
    choice. It tests whether the same clones rank consistently differently
    across the four planting sites.
  >> VARIABLE SELECTION:
    - Repeated: site_Antalya
    - Repeated: site_Ankara
    - Repeated: site_Eskisehir
    - Repeated: site_Konya

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Friedman chi-square(3) = 11.88   p = 0.008 **   Kendall W = 0.198   n = 20
    The same clones' growth across 4 sites differs significantly.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same clones' growth across four trial sites (Antalya, Ankara, Eskisehir, Konya) with the Friedman
    test -- the nonparametric counterpart of repeated-measures ANOVA, and effectively a G×E test. The result is
    significant (chi-square(3) = 11.88, p = 0.008), with Kendall W = 0.20 indicating low consistency of the ranking
    across sites -- this is important: the same clone ranks differently at different sites, i.e. there is strong
    Genotype×Environment interaction. Because the same clones appear at all sites, observations are dependent; Friedman
    accounts for this correctly. It shows that clonal forestry needs site-specific clone selection.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 018_binomial_PopulusNigra_sex_ratio_vs_50_50.xlsx
  >> SCENARIO (narration):
    Populus nigra is dioecious, and we want to know whether the observed sex
    ratio departs from the expected balanced 50:50. Across several riparian
    populations including Goksu_delta, Goksun_valley, sakarya_Adapazari,
    yesilirmak_valley, Saros_gulf and Ceyhan_valley, each tree was scored as
    sexual_male_1_female_0. Since the outcome is binary with a clear theoretical
    proportion of 0.5, a Binomial Test directly evaluates whether the proportion
    of males differs significantly from one half. This tells us whether the
    species shows a genuine sex-ratio bias in these stands.
  >> VARIABLE SELECTION:
    - Outcome: sexual_male_1_female_0
    - Group: population

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed male proportion = 0.6136   test proportion = 0.50
    p = 0.0009 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the sex ratio in a Populus nigra (black poplar -- a dioecious, separate-sexed species)
    population departs from 50:50 with the binomial test. The observed male proportion is 0.61 -- significantly above
    the expected 50% (p < .001). So males dominate this population. Sex-ratio skew matters for population genetics and
    reproductive biology: an unbalanced sex ratio lowers the effective population size (Ne) and, over the long term,
    threatens genetic diversity. The binomial test is the simplest way to compare a single binary proportion
    (male/female) against a known expectation (50:50).

====================================================================================

#19  Sign Test
    file: 019_sign_test_bud_median_climate_kaymasi.xlsx
  >> SCENARIO (narration):
    To detect a shift in the median spring phenology of nurseries over two
    decades, we recorded budburst date as bud_jd_2000 and bud_jd_2020 in Julian
    days for the same nurseries. We are interested only in the direction of
    change within each nursery, not its magnitude, and we make no assumptions
    about the shape of the difference distribution. The Sign Test simply counts
    how many nurseries advanced versus delayed and tests whether the median
    change differs from zero. This provides a minimal-assumption check for a
    climate-driven shift in budburst timing.
  >> VARIABLE SELECTION:
    - Measurement 1: bud_jd_2000
    - Measurement 2: bud_jd_2020

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive (2000 > 2020) = 35   (bud day mostly shifted EARLIER)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We compared the bud-flush day (Julian day) in the same nurseries in 2000 and 2020 with the paired sign test. In
    most nurseries (35) the 2000 day was higher than the 2020 day -- i.e. bud flush shifted to an earlier day in 2020
    (p < .001). This is strong evidence that climate warming pulls tree phenology earlier: spring arrives sooner,
    trees wake earlier. The sign test looks only at the direction of change (earlier/later); it is a robust,
    assumption-free test. Phenological shift is one of the clearest biological fingerprints of climate change.

====================================================================================

#20  Runs Test
    file: 020_runs_test_bud_induction_daily.xlsx
  >> SCENARIO (narration):
    We monitored new_bud_induction_count on a daily basis to examine whether bud
    induction events occur in a random temporal sequence or are clustered in
    runs. Each successive day_no carries a count, and we want to know whether
    high and low induction days alternate randomly or form non-random streaks.
    The Runs Test evaluates the ordering of the daily series for departures from
    randomness. Detecting clustering would suggest that environmental triggers
    drive bud induction in bursts rather than independently from day to day.
  >> VARIABLE SELECTION:
    - Sequence: new_bud_induction_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 58   Z = -0.11   (accepted as random)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the daily bud-induction count series is randomly distributed -- whether successive days cluster
    or show a pattern -- with the Wald-Wolfowitz runs test. With Z = -0.11 the observed number of runs is very close
    to expected; the series is accepted as random, with no systematic pattern. This matters methodologically: if a
    series is non-random (autocorrelated), many classic tests become invalid. Because randomness holds here, standard
    methods can be used safely. The runs test is ideal for this pre-check on time-ordered phenology/growth data.

====================================================================================

#21  Chi-Square Test of Independence
    file: 021_chisquare_independence_rare_allele_population.xlsx
  >> SCENARIO (narration):
    In this study we ask whether a rare allele at the SSR_07 locus is
    distributed independently of the seed source population across four
    Anatolian provenances. We screened 240 trees and recorded, for each, the
    population they came from (kizilirmak, sakarya, Coruh_valley, Goksu) and
    whether the rare_allele_present marker was detected. Because both variables
    are categorical and we want to know if presence of the rare allele depends
    on provenance, a Chi-Square Test of Independence is the appropriate choice.
    A significant result would indicate that the rare allele is concentrated in
    certain seed sources rather than spread evenly, which matters for
    conservation prioritisation.
  >> VARIABLE SELECTION:
    - Row variable: population
    - Column variable: rare_allele_present

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 4.7028   p = 0.195 ns   Cramer V = 0.14 (Weak)
    NO significant association between rare-allele presence and population.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether rare-allele presence is associated with population using chi-square. The result is
    non-significant (chi-square(3) = 4.70, p = 0.195, Cramer V = 0.14): rare alleles are not distributed differently
    among populations -- they appear at similar frequency regardless of population. A non-significant result is
    information too: rare alleles are not specific to a particular population but spread across the range (which may
    point to gene flow distributing them, or shared ancestral polymorphism). For association between two categorical
    variables the chi-square test of independence is the standard method.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 022_chisquare_goodnessfit_HW_dengesi_genotype_frequency.xlsx
  >> SCENARIO (narration):
    Here we test whether the genotype frequencies in our forest stand match the
    expectations of Hardy-Weinberg equilibrium. We genotyped 360 individuals and
    classified each at a single locus as AA, Aa, or aa. To compare these
    observed counts against the theoretical Hardy-Weinberg proportions, a Chi-
    Square Goodness-of-Fit test is used. A significant deviation would suggest
    forces such as inbreeding, selection, or non-random mating are disturbing
    the population's genetic balance.
  >> VARIABLE SELECTION:
    - Category variable: genotype

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(2) = 64.0667   p < .001 ***   n = 360   k = 3 genotypes (expected: HW ratios)
    H0 REJECTED (deviation from Hardy-Weinberg equilibrium)

>> COMMENTARY (narration):
    We tested whether the genotype frequencies at a locus (AA, Aa, aa) fit the proportions expected under Hardy-
    Weinberg (HW) equilibrium with the chi-square goodness-of-fit test. The observed distribution deviates very
    significantly from the HW expectation (chi-square(2) = 64.07, p < .001). This is the most fundamental test in
    population genetics: deviation from HW says something is "abnormal" in the population -- inbreeding, selection,
    gene flow, small population, or the Wahlund effect. The direction of the deviation (heterozygote deficit/excess)
    hints which process is at work. The goodness-of-fit test is the right way to compare the observed genotype
    distribution against the theoretical HW expectation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 023_fisher_exact_QTL_haplotype_dbh_small_N.xlsx
  >> SCENARIO (narration):
    This study explores whether carrying a particular QTL haplotype is
    associated with superior diameter growth in a small breeding trial. For each
    of just 25 saplings we recorded the genotype_QTL_haplo_X status (present or
    absent) and whether the tree expressed a high_dbh_phenotype. Because the
    sample size is small and some expected cell counts will be low, Fisher's
    Exact Test is the correct way to evaluate this 2x2 association. A
    significant result would support using the haplotype as a marker for
    selecting fast-growing genotypes.
  >> VARIABLE SELECTION:
    - Group variable: genotype_QTL_haplo_X
    - Outcome variable: high_dbh_phenotype

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher Exact p = 0.0154 *   Odds Ratio = 11.25  (95% CI: 1.65 - 76.85)   Phi = 0.53 (Strong)
    QTL haplotype is associated with the high-diameter phenotype.   H0 REJECTED

>> COMMENTARY (narration):
    We tested the association between a QTL (quantitative trait locus) haplotype and the high stem-diameter (dbh)
    phenotype with Fisher's exact test -- because chi-square is unreliable for tables with few observations (small
    cells). The result is significant (p = 0.015), odds ratio 11.25: trees carrying this haplotype are 11 times more
    likely to be high-diameter. Phi = 0.53 indicates a strong association. This is very valuable for marker-assisted
    selection (MAS): the haplotype may be a genetic marker of the superior phenotype. In small-sample QTL/association
    studies Fisher's exact test gives more accurate results than chi-square; it is often needed in genetic
    association scans.

====================================================================================

#24  McNemar Test
    file: 024_mcnemar_marker_asset_yillara_gore.xlsx
  >> SCENARIO (narration):
    We examine whether the presence of a diagnostic genetic marker in the same
    trees changed between two assessment years. The identical 50 individuals
    were scored for marker_asset_2010 and again for marker_asset_2024, giving
    paired binary observations on each tree. Since the measurements are paired
    and binary, the McNemar Test is used to detect a systematic shift in marker
    presence over time. A significant change would indicate the marker's
    detectability or frequency moved across the 14-year interval rather than
    staying stable.
  >> VARIABLE SELECTION:
    - Before measurement: marker_asset_2010
    - After measurement: marker_asset_2024

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar (exact binomial) min(b,c) = 3   p = 0.146 ns   discordant: b = 9, c = 3
    H0 NOT REJECTED (no significant change)

>> COMMENTARY (narration):
    We compared the presence of a marker in the same individuals in 2010 and 2024 -- paired binary data call for
    McNemar's test (in the context of genotyping reproducibility/technology change). The result is non-significant
    (p = 0.146): marker status changed one way in 9 individuals and the other way in 3 -- the balance is not tipped
    enough, so there is no significant shift in overall marker presence. The small number of discordant cells limits
    the test's power. A non-significant result is information too: there is no systematic difference in marker
    detection between the two dates (genotyping is consistent). McNemar looks only at the cells that changed; it is
    the right way to measure paired binary change.

====================================================================================

#25  Cohen's Kappa
    file: 025_cohen_kappa_morphological_Quercus_taksonom.xlsx
  >> SCENARIO (narration):
    This study evaluates how consistently two taxonomists assign oak species
    based on morphological traits. Seventy samples were independently classified
    by expert_A_morphological_diagnosis and expert_B_morphological_diagnosis
    into one of four Quercus species (Q_petraea, Q_cerris, Q_pubescens,
    Q_robur). To quantify agreement beyond what chance alone would produce, we
    compute Cohen's Kappa on these paired categorical ratings. A high kappa
    would confirm that morphological diagnosis is reproducible before it is used
    alongside genetic markers.
  >> VARIABLE SELECTION:
    - Rater 1: expert_A_morphological_diagnosis
    - Rater 2: expert_B_morphological_diagnosis

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's Kappa = 0.9237 (almost perfect agreement)   p < .001 ***
    Two experts' morphological Quercus identification

>> COMMENTARY (narration):
    We measured the agreement between two experts independently identifying Quercus (oak) samples morphologically with
    Cohen's Kappa. Kappa = 0.92 is in the "almost perfect" range -- a chance-corrected measure of agreement, i.e. the
    real consistency after removing accidental agreement. In plant taxonomy this matters greatly: oak species are
    morphologically hard to tell apart (widespread hybridisation), so if experts identified them inconsistently all
    species records would be suspect. Kappa = 0.92 shows the morphological identifications are very reliable and
    reproducible. Reporting inter-rater reliability is the quality assurance of taxonomic studies.

====================================================================================

#26  Cochran-Mantel-Haenszel Test
    file: 026_cmh_sex_strata_marker_phenotype.xlsx
  >> SCENARIO (narration):
    We test whether the presence of a genetic marker is associated with a high
    diameter phenotype while controlling for sex as a stratifying variable.
    Across 240 trees we recorded sex (male, female), marker_status
    (marker_present, marker_absent), and whether each showed a
    high_dbh_phenotype. The Cochran-Mantel-Haenszel Test lets us assess the
    marker-phenotype association pooled over the sex strata, so any confounding
    by sex is accounted for. A significant common odds ratio would support a
    real, sex-independent link between the marker and superior growth.
  >> VARIABLE SELECTION:
    - Stratum: sex
    - Row variable: marker_status
    - Column variable: high_dbh_phenotype

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi-square(1) = 21.34   p < .001 ***   Common OR (MH) = 3.57  (95% CI: 2.06 - 6.18)
    Breslow-Day chi-square(1) = 1.46, p = 0.23 (homogeneous OR)   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between marker status and the high-diameter phenotype while controlling for sex
    (strata) with the CMH test. Even holding sex constant the association is strong: common odds ratio 3.57 (p < .001)
    -- stripped of the sex confounder, marker carriers are 3.6 times more likely to be high-diameter. The Breslow-Day
    test is non-significant (p = 0.23), meaning the association is consistent in both sexes (homogeneous OR). This
    confirms the marker is a sex-independent, genuine genetic indicator -- added confidence for marker-assisted
    selection. CMH is the classic way to control for a third variable (sex) by stratification.

====================================================================================

#27  Log-Linear Analysis
    file: 027_log_lineer_type_season_reproduction.xlsx
  >> SCENARIO (narration):
    This study models how reproductive organ counts vary jointly with pine
    species and season. We have aggregated frequency counts cross-classified by
    type (P_brutia, P_nigra), season (spring, summer), and reproduction_organi
    (male_cone, female_cone). A Log-Linear Analysis is used to examine the
    associations and interactions among these three categorical factors using
    the cell frequencies. The model reveals, for example, whether the balance of
    male versus female cones depends on species, season, or their combination.
  >> VARIABLE SELECTION:
    - Factor: type
    - Factor: season
    - Factor: reproduction_organi
    - Count: frequency

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Best model AIC = 50.91   Pearson chi-square = 0.08 (model fits the data very well)
    Factors: type x season x reproductive organ

>> COMMENTARY (narration):
    We analysed the contingency structure formed jointly by three categorical variables -- type, season and
    reproductive organ -- with a log-linear model. The selected model fits very well (Pearson chi-square = 0.08, very
    low), with AIC = 50.91 as the most parsimonious model. Log-linear analysis lets us untangle interactions among
    more than two categorical variables (which are dependent, which interaction is significant) -- it is like ANOVA
    for categorical data. It is powerful for summarising the multi-way relationships of reproductive phenology (which
    type, in which season, with which reproductive organ); used in reproductive biology and phenology studies.

====================================================================================

#28  Crosstabulation
    file: 028_cross_tablo_clone_region_distribution.xlsx
  >> SCENARIO (narration):
    Here we describe how four selected clones are distributed across two
    geographic regions of a clonal seed orchard. For 200 ramets we recorded the
    clone identity (clone_1 through clone_4) and the region in which it was
    planted (Marmara, black_sea). A Crosstabulation summarises the joint counts
    of clone by region in a contingency table. This overview shows whether
    certain clones dominate particular regions, informing the spatial balance of
    the orchard.
  >> VARIABLE SELECTION:
    - Row variable: clone
    - Column variable: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 11.84   p = 0.008 **   Cramer V = 0.24 (Moderate)
    Clone is associated with region distribution.   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between clone and region with a crosstab and chi-square. The association is significant
    and moderate (chi-square(3) = 11.84, p = 0.008, Cramer V = 0.24): the distribution of clones across regions is not
    random -- some clones concentrate in particular regions. In clonal forestry this shows a clone-region pattern
    (which clone is planted where), which may stem from local adaptation or breeding/planting preferences. The crosstab
    is the most basic and readable way to see the co-occurrence of two categorical variables; Cramer V measures the
    strength.

====================================================================================

#29  MR Frequency
    file: 029_MR_frequency_breeding_target_secimleri.xlsx
  >> SCENARIO (narration):
    In this survey we summarise which breeding objectives forestry professionals
    prioritise. One hundred participants reported their institution (TARGEM,
    ogm, university, private_Sirket), their experience_year, and selected their
    breeding_target_oncelikleri, a multiple-response item where each respondent
    could choose several priorities. A Multiple Response Frequency analysis
    tabulates how often each breeding target was chosen across all selections.
    The resulting frequencies reveal the most commonly endorsed breeding goals
    in the sector.
  >> VARIABLE SELECTION:
    - Multiple response set: breeding_target_oncelikleri

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 100   Respondents = 100 (100%)
    Breeding targets (growth rate, wood quality, drought resistance, disease resistance...) selected as multi-choice

>> COMMENTARY (narration):
    Forest geneticists/breeders were asked which targets they prioritise in breeding programs -- a "multiple response"
    question where several options can be ticked. All 100 experts selected at least one target, so the response rate
    is 100%. Multiple-response frequency analysis shows how many times and by what percentage of experts each target
    was chosen; the totals exceed 100% because everyone can tick several boxes. Profiling the distribution of breeding
    priorities (growth, quality, resistance) in expert opinion is a valuable tool in setting national breeding strategy.

====================================================================================

#30  MR Crosstab
    file: 030_MR_categorical_genetic_threats_institution.xlsx
  >> SCENARIO (narration):
    This survey examines how perceptions of genetic threats to forest
    populations differ by the type of institution experts work for. Among 130
    experts we recorded their institution (ogm, university, AGM) and their
    perceived_genetic_threats, a multiple-response item allowing each expert to
    flag several threats. A Multiple Response Crosstab pairs the multiple-
    response threat set against institution to show how threat awareness varies
    across organisations. The table highlights which threats are emphasised by
    which institutional group.
  >> VARIABLE SELECTION:
    - Row variable: institution
    - Multiple response set: perceived_genetic_threats

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 130   Group column: institution
    Perceived genetic threats crossed by institution

>> COMMENTARY (narration):
    We cross-tabulated the genetic threats experts perceive (multiple-choice -- genetic erosion, habitat loss, climate,
    inbreeding...) by their institution. With data from 130 experts we can see how each threat perception is
    distributed across institutions. The multiple-response crosstab answers "which institution stresses which threats
    more" -- e.g. a university vs an applied agency may show different priorities. This reveals inter-institution
    differences in perception of forest-genetic-resource conservation and guides the development of common policy.

====================================================================================

#31  MR x MR Crosstab
    file: 031_MR_MR_target_tehdit.xlsx
  >> SCENARIO (narration):
    In a survey of forest geneticists and seed-orchard managers, we asked each
    participant about their primary breeding objectives and about the main
    threats they perceive to seed sources. We want to see whether the chosen
    hedefler (breeding goals) are associated with the kinds of threats reported,
    so that conservation priorities can be aligned with stakeholder concerns.
    Because both hedefler and threats are categorical responses, a contingency
    crosstab with a chi-square test of independence is the right tool to reveal
    which goal-threat combinations occur more often than chance.
  >> VARIABLE SELECTION:
    - Row variable: hedefler
    - Column variable: threats

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 80
    Breeding targets x perceived threats (habitat, climate, disease, inbreeding) cross-tabulated

>> COMMENTARY (narration):
    This is the most advanced multiple-response analysis: we cross two separate multi-choice questions against each
    other -- the expert's breeding targets and perceived genetic threats. Among 80 experts we can see which targets-
    prioritisers stress which threats: e.g. those prioritising drought resistance may stress the climate threat, while
    those prioritising diversity conservation stress inbreeding/habitat loss. This table reveals target-threat
    pairings, providing powerful data for forest genetic-resource management strategy.

====================================================================================

#32  Cochran's Q Test
    file: 032_cochran_Q_seedling_survival_survival_4yil.xlsx
  >> SCENARIO (narration):
    We monitored the same set of seedlings in a provenance trial and recorded
    whether each one was alive at the end of year1, year2, year3, and year4. The
    research question is whether cumulative survival changes significantly
    across these four annual assessments. Since the survival outcome is binary
    (1/0) and measured repeatedly on the identical seedlings, Cochran's Q test
    is the appropriate way to test for a difference in survival proportions over
    the four years.
  >> VARIABLE SELECTION:
    - Repeated measure: year1
    - Repeated measure: year2
    - Repeated measure: year3
    - Repeated measure: year4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 44.64   p < .001 ***
    Seedling survival (binary) changes significantly over 4 years.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same seedlings' survival status (alive/dead binary) over four years with Cochran's Q test -- the
    binary counterpart of repeated measures for more than two occasions, the multi-occasion extension of McNemar. The
    result is highly significant (Q = 44.64, p < .001): survival changes markedly across years (it declines) --
    seedlings have annual mortality. This is critical in monitoring afforestation success: early-year mortality
    reflects planting success and the suitability of provenance/growing conditions. Cochran's Q is the right way to
    test change in repeated binary (survival) measurements on the same units.

====================================================================================

#33  Correlation Analysis
    file: 033_correlation_kantitatif_ozellikler_PinusBrutia.xlsx
  >> SCENARIO (narration):
    For a sample of Pinus brutia trees we measured several quantitative traits:
    age_year, dbh_cm, height_m, wood specific gravity (wsg_gcm3), and
    branch_count. We want to understand how these growth and wood-quality
    variables move together, for example whether diameter and height are tightly
    coupled or whether wood density trades off against fast growth. A
    correlation analysis quantifies the strength and direction of the pairwise
    linear relationships among all these continuous traits.
  >> VARIABLE SELECTION:
    - Variable: age_year
    - Variable: dbh_cm
    - Variable: height_m
    - Variable: wsg_gcm3
    - Variable: branch_count

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Strongest relationship: age - height   r = +0.869   p < .001 ***  (Strong)
    5 variables: age, dbh, height, wood density, branch count

>> COMMENTARY (narration):
    We computed all pairwise correlations among five quantitative tree traits. The strongest is between age and height
    (r = 0.87) -- an expected result, older trees are taller. Generally height and diameter are highly correlated,
    while wood density is weakly related to them (this matters for breeding: it may be possible to select for fast
    growth and high wood density together, or a trade-off may be needed). The correlation matrix is the first step
    before building a regression/selection index -- to see the structure among traits and detect collinearity. In
    multi-trait tree breeding, genetic/phenotypic correlations among traits shape the selection strategy.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Analysis
    file: 034_bland_altman_RAPD_SSR_He_goodnessfit.xlsx
  >> SCENARIO (narration):
    At each genetic locus we estimated expected heterozygosity (He) twice, once
    from RAPD markers (RAPD_He) and once from SSR markers (SSR_He), to check
    whether the two marker systems give interchangeable diversity estimates. The
    goal is not merely to correlate them but to assess agreement and any
    systematic bias between the methods. A Bland-Altman analysis is ideal here
    because it plots the difference against the mean of RAPD_He and SSR_He,
    revealing bias and limits of agreement across loci.
  >> VARIABLE SELECTION:
    - Method 1: RAPD_He
    - Method 2: SSR_He

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    RAPD He mean = 0.313   |   SSR He mean = 0.295   |   Bias ~ +0.018 (methods agree largely)

>> COMMENTARY (narration):
    We compared two molecular marker methods of measuring genetic diversity (He) -- RAPD and SSR (microsatellite) --
    on the same loci with Bland-Altman. This method does not look at correlation (high correlation does not guarantee
    agreement); it measures the difference between the two measurements. The means are close (0.313 vs 0.295), the
    bias about 0.018 -- the two marker systems estimate He almost the same, giving comparable results (RAPD slightly
    higher). Bland-Altman is the standard way to assess the agreement of two molecular methods (old/new marker, two
    platforms); it is used when switching/comparing methods in genetic-diversity studies.

====================================================================================

#35  Effect Size
    file: 035_effect_size_marginal_center_He.xlsx
  >> SCENARIO (narration):
    We compared expected heterozygosity between marginal and center populations
    to test the common hypothesis that range-edge stands are genetically
    depauperate. The population_type factor distinguishes marginal versus center
    populations, and mean_He records the genetic diversity for each. Beyond
    simple significance, we want to quantify how large the diversity gap is, so
    an effect size analysis expresses the magnitude of the He difference between
    the two population types in standardized terms.
  >> VARIABLE SELECTION:
    - Group: population_type
    - Measure: mean_He

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = 0.93  (large effect)
    He difference between marginal vs central populations

>> COMMENTARY (narration):
    We measured the heterozygosity (He) difference between marginal (range-edge) and central populations not just as
    "is it significant" but "how large is it" -- as an effect size. Cohen's d = 0.93 is a large effect; the two groups
    clearly separate (central populations are more diverse). This supports the population-genetics "central-marginal
    hypothesis": populations at the range centre carry more genetic diversity, while edge populations are poorer due to
    bottlenecks/isolation. A large effect of d = 0.93 has practical importance for setting conservation priorities
    (which populations are more valuable). Reporting effect size is standard in publications.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 036_cca_environment_allele_frekansi_relationship.xlsx
  >> SCENARIO (narration):
    For a set of populations we recorded environmental descriptors (lat, lon,
    elevation_m, precipitation_mm, temperature_C) alongside the frequencies of
    four alleles (freq_allele_A1, freq_allele_A2, freq_allele_B1,
    freq_allele_B2). The research question is whether the multivariate
    environmental gradient is jointly related to the multivariate pattern of
    allele frequencies, which would indicate environmental selection shaping
    genetic structure. Canonical correlation analysis is suited to this because
    it finds the linear combinations of the environmental set and the allele-
    frequency set that are maximally correlated.
  >> VARIABLE SELECTION:
    - Set X: lat
    - Set X: lon
    - Set X: elevation_m
    - Set X: precipitation_mm
    - Set X: temperature_C
    - Set Y: freq_allele_A1
    - Set Y: freq_allele_A2
    - Set Y: freq_allele_B1
    - Set Y: freq_allele_B2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.925, chi-square(20) = 276.07, p < .001   |   CC2: r = 0.903, p < .001
    X set: environment (lat/lon/elevation/precipitation/temperature)  |  Y set: allele frequencies (A1/A2/B1/B2)

>> COMMENTARY (narration):
    We related two multivariate sets -- environmental variables (location, elevation, precipitation, temperature) and
    allele frequencies -- with canonical correlation. The first canonical axis is very strong (r = 0.93, p < .001),
    the second also significant (r = 0.90): environment and genetic structure are tightly linked. This is an important
    pattern in population genetics -- allele frequencies change along an environmental gradient, i.e. evidence of
    environment-related selection (local adaptation) or isolation-by-environment. CCA is the most direct answer to
    "how do two blocks of variables co-vary"; it is a powerful method in landscape genetics for studying the
    environment-gene relationship.

====================================================================================

#37  Correspondence Analysis (CA)
    file: 037_correspondence_clone_region_kontingenz.xlsx
  >> SCENARIO (narration):
    We tallied how often each clone appears across five geographic regions
    (mediterranean, Marmara, anatolia, black_sea, EgeIc) to explore the regional
    distribution of clonal material in our breeding program. The question is
    which clones are characteristic of which regions, revealing patterns of
    provenance and deployment. Correspondence analysis is appropriate for this
    large region-by-clone contingency table, mapping rows and columns into a
    shared low-dimensional space that exposes their associations.
  >> VARIABLE SELECTION:
    - Row variable: region
    - Column variable: clone

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.307   Number of dimensions = 2
    Clone x region contingency table

>> COMMENTARY (narration):
    We projected the contingency table of clone by region into a two-dimensional map with correspondence analysis.
    Total inertia 0.307 indicates a moderate-strong association; the two dimensions visualise most of that structure.
    Clone-region pairs that fall close on the map indicate that the clone is associated with that region. In clonal
    forestry this lets us read at a glance which clones concentrate in which regions -- clone-region adaptation or
    planting pattern. Correspondence analysis is a powerful way to visualise categorical relationships and gives the
    association strength (inertia) numerically.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 038_varclus_15kantitatif_feature_clustering.xlsx
  >> SCENARIO (narration):
    Across many trees we measured fifteen quantitative features spanning growth
    over time (height_year1 through height_year4, diameter_year3), stress
    responses (drought_score, cold_score, disease_score), survival, and wood and
    architectural traits (wsg, MFA_mikrofibril, branch_angle, branch_count,
    leaf_angle, age_half). With so many partly redundant variables, we want to
    group them into clusters of related traits to simplify later modeling and
    selection indices. Variable clustering (VarClus) is the right approach
    because it partitions the feature set into clusters of mutually correlated
    variables.
  >> VARIABLE SELECTION:
    - Variable: height_year1
    - Variable: height_year2
    - Variable: height_year3
    - Variable: height_year4
    - Variable: diameter_year3
    - Variable: drought_score
    - Variable: cold_score
    - Variable: disease_score
    - Variable: survival
    - Variable: wsg
    - Variable: MFA_mikrofibril
    - Variable: branch_angle
    - Variable: branch_count
    - Variable: leaf_angle
    - Variable: age_half

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (15 quantitative tree traits)
    A representative trait was selected from each cluster

>> COMMENTARY (narration):
    We clustered 15 quantitative tree traits (annual heights, diameter, drought/cold/disease scores, wood density,
    branch angle, etc.) by similarity into three groups. VarClus clusters traits, not observations: it shows which
    traits measure the same "dimension". The three clusters probably represent growth (heights/diameter), stress
    resistance (drought/cold/disease) and wood quality. This lets us pick one representative trait per group and
    reduce many measurements -- extremely practical in tree breeding when building selection indices and reducing
    phenotyping effort.

====================================================================================

#39  Multiple Linear Regression
    file: 039_multiple_linear_regression_height_env_BV.xlsx
  >> SCENARIO (narration):
    We want to predict five-year height growth (height_cm_5yr) of young trees
    from their planting-site environment (elevation_m, precipitation_mm,
    temperature_C) together with their genetic potential captured by the
    parent's breeding value (parent_BLUP). The aim is to separate environmental
    drivers from inherited growth potential so that site matching and parent
    selection can be optimized. Multiple linear regression fits height_cm_5yr as
    a function of these continuous predictors, estimating each one's independent
    contribution.
  >> VARIABLE SELECTION:
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Predictor: temperature_C
    - Predictor: parent_BLUP
    - Dependent: height_cm_5yr

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.890   Adjusted R2 = 0.888
    Height ~ elevation + precipitation + temperature + parent BLUP

>> COMMENTARY (narration):
    We built a multiple linear regression predicting 5-year height from four variables -- elevation, precipitation,
    temperature and parent breeding value (BLUP). The model is very strong: together they explain 89% of height
    variance (R2 = 0.89). This is a critical forest-genetics result: height is strongly determined by both
    environmental conditions and the parent's genetic value (BLUP) -- i.e. both "where it is planted" and "its
    genetics" matter. BLUP being strong in the model shows genetic gain can be achieved through parent selection.
    Multiple regression is the basic tool for explaining a continuous trait with environmental and genetic drivers
    together.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 040_logistic_5yr_survival_survival_site_provenance.xlsx
  >> SCENARIO (narration):
    We recorded whether each seedling survived to five years
    (5yr_survival_survival_1_0) and want to know how survival probability
    depends on the seedling's provenance (Mugla_Marmaris, Burdur_Aglasun,
    Mersin_Erdemli, Mut_Karaman, Kahramanmaras_Andirin, Adana_Pos) and on the
    planting site's elevation (site_elevation_m) and temperature
    (site_temperature_C). The purpose is to identify which seed sources and site
    conditions favor establishment success. Because the outcome is binary,
    logistic regression is the appropriate model to estimate how provenance and
    site climate shape the odds of survival.
  >> VARIABLE SELECTION:
    - Predictor: provenance
    - Predictor: site_elevation_m
    - Predictor: site_temperature_C
    - Dependent: 5yr_survival_survival_1_0

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: Logistic Regression   Pseudo R2 = 0.118
    5-year survival (1/0) ~ site elevation + site temperature

>> COMMENTARY (narration):
    We built a logistic regression predicting whether a seedling survives to 5 years from site elevation and
    temperature. Because the outcome is binary (alive/dead), logistic regression is the right choice. Pseudo R2 =
    0.118 shows the model partly explains survival -- the coefficients generally say survival rises under certain
    elevation/temperature conditions. This is valuable for afforestation site selection and provenance-site matching
    (assisted migration): it shows under which environmental conditions seedlings survive more. In forest genetics,
    logistic regression is widely used for survival/adaptation prediction.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Poisson Regression
    file: 041_poisson_branch_count_age_site.xlsx
  >> SCENARIO (narration):
    In a forest provenance trial we want to understand what drives branching
    architecture in young trees, since side branch counts affect both timber
    quality and crown development. We recorded the side_branch_count for each
    tree along with its age_year and the planting site (sites A, B, and C).
    Because side_branch_count is a count of discrete events that cannot be
    negative, Poisson Regression is the natural model to relate this count
    outcome to tree age and site. This lets us quantify how branch production
    increases with age and whether site conditions shift the branching rate.
  >> VARIABLE SELECTION:
    - Outcome (count): side_branch_count
    - Predictor: age_year
    - Predictor (categorical): site

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 1321.10   Deviance = 362.69
    Side-branch count ~ age   (Poisson)

>> COMMENTARY (narration):
    We built a Poisson regression predicting trees' side-branch count -- a count variable -- from age. Count data
    (0,1,2,...) are not normally distributed, so we use Poisson rather than linear regression. The model typically
    shows branch count rising with age. Deviance and AIC assess model fit. Branch count is an important breeding trait
    for stem form/wood quality (fewer branches = straighter stem). In forest genetics, count regression is the correct
    way to relate count outcomes (branch count, cone count, seedling count) to predictors.

====================================================================================

#42  Multinomial Logistic Regression
    file: 042_multinomial_3sinif_cold_zarari.xlsx
  >> SCENARIO (narration):
    After an unexpected cold spell hit our seedling nurseries, we classified
    each seedling's cold_damage_class as low, medium, or high to understand
    provenance susceptibility to frost. We want to predict this three-category
    damage outcome from the seedling's provenance (six seed sources such as
    Mut_Karaman and Antalya_Duzlercami), the site_elevation_m, and the recorded
    min_temperature_C. Because cold_damage_class has three unordered nominal
    categories, Multinomial Logistic Regression is the appropriate model. The
    results will reveal which provenances and environmental conditions raise the
    odds of severe cold damage.
  >> VARIABLE SELECTION:
    - Outcome (nominal): cold_damage_class
    - Predictor (categorical): provenance
    - Predictor: site_elevation_m
    - Predictor: min_temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    3 cold-damage classes (low/medium/high)   Reference class: 'high'
    Cold-damage class ~ site elevation + min temperature

>> COMMENTARY (narration):
    We split cold damage in seedlings into three classes (low/medium/high) and modelled this class from site elevation
    and minimum temperature with multinomial logistic regression. Because there are more than two unordered categories,
    this method fits: 'high damage' is the reference, and a separate equation is built for each other class.
    Coefficients read as "how the odds of being in a given damage class change as elevation/temperature shift". This
    relates frost-damage risk to environmental conditions -- valuable for provenance-site matching and frost-risk
    mapping. It is used to assess seedling adaptation under climate change's increasing late-frost events.

====================================================================================

#43  Ordinal Logistic Regression
    file: 043_ordinal_5sinif_cold_damage_score.xlsx
  >> SCENARIO (narration):
    To evaluate frost hardiness across seed sources, we scored each seedling's
    cold damage on an ordered five-point scale, cold_damage_5sinif, ranging from
    1_low to 5_high. We aim to predict this ordered damage score from the
    seedling's provenance (six sources including Burdur_Aglasun and
    Mugla_Marmaris) and the elevation_m of its origin. Since the outcome is an
    ordered category where higher classes mean progressively worse damage,
    Ordinal Logistic Regression respects this natural ranking. This helps us
    rank provenances by their cold tolerance while accounting for elevation of
    origin.
  >> VARIABLE SELECTION:
    - Outcome (ordinal): cold_damage_5sinif
    - Predictor (categorical): provenance
    - Predictor: elevation_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 414.56
    Cold-damage score (1-5, ordered) ~ site elevation

>> COMMENTARY (narration):
    We built an ordinal logistic regression predicting cold-damage score (1-5, ordered class) from site elevation.
    These classes are ordered (1=undamaged ... 5=heavily damaged), and the ordinal model uses that order, making it
    more powerful and interpretable than multinomial. The coefficients give the tendency to move to a higher damage
    class as elevation rises (frost risk increases at high altitude). For ordered categorical outcomes (damage score,
    disease score, quality tier), ordinal logistic is the right method; it is widely used in forest genetics for
    damage/score assessments.

====================================================================================

#44  PLS Regression
    file: 044_pls_8marker_correlated_BV_height.xlsx
  >> SCENARIO (narration):
    In a genomic selection study we have eight molecular markers,
    marker_1_correlated_BV through marker_8_correlated_BV, that are highly
    correlated with one another because they tag overlapping breeding-value
    signals. We want to predict early growth, height_cm_3yr, from these
    correlated markers to support early selection of superior trees. Because the
    eight predictors are strongly collinear, ordinary regression is unstable, so
    PLS Regression is ideal as it extracts latent components that capture the
    shared marker variation. This gives a robust predictive model linking marker
    profiles to three-year height.
  >> VARIABLE SELECTION:
    - Predictor: marker_1_correlated_BV
    - Predictor: marker_2_correlated_BV
    - Predictor: marker_3_correlated_BV
    - Predictor: marker_4_correlated_BV
    - Predictor: marker_5_correlated_BV
    - Predictor: marker_6_correlated_BV
    - Predictor: marker_7_correlated_BV
    - Predictor: marker_8_correlated_BV
    - Outcome: height_cm_3yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2_train = 0.910   R2_CV (5-fold) = 0.874
    3-year height ~ 8 markers (correlated with BV)

>> COMMENTARY (narration):
    We predicted 3-year height from 8 molecular markers correlated with the breeding value (BV) with PLS (Partial
    Least Squares) regression. When markers are correlated, ordinary regression becomes unstable; PLS solves this by
    reducing the markers to a few orthogonal components. The model is very strong both in training (R2 = 0.91) and
    cross-validation (R2_CV = 0.87) -- no overfitting, good generalisation. This is the logic of genomic selection:
    predicting the phenotype/breeding value from many markers. For high-dimensional, collinear marker data, PLS keeps
    both predictive power and interpretability.

====================================================================================

#45  Probit Regression
    file: 045_probit_bud_fracture_chill.xlsx
  >> SCENARIO (narration):
    Studying spring frost risk, we recorded for each seedling whether a bud
    fractured under chilling, coded as bud_fracture_1_0. We want to model the
    probability of bud fracture as a function of accumulated chill_hour and
    seedling size measured by dbh_cm. Because the outcome is binary and we
    assume a normally distributed latent susceptibility threshold, Probit
    Regression is a fitting choice. This quantifies how chilling exposure and
    stem diameter shift the likelihood of bud damage.
  >> VARIABLE SELECTION:
    - Outcome (binary): bud_fracture_1_0
    - Predictor: chill_hour
    - Predictor: dbh_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 236.60   Pseudo R2 (McFadden) = 0.166
    Bud flush (1/0) ~ chill hours + dbh

>> COMMENTARY (narration):
    We predicted whether a seedling flushes its buds from chill hours and stem diameter with probit regression. Probit
    models a binary outcome like logistic but assumes a normal-distribution curve (cumulative normal). Pseudo R2 =
    0.166: chill hours partly explain bud flush -- seedlings that accumulate enough chilling flush more (the basic
    mechanism of plant phenology: breaking winter dormancy requires chilling). This is critical for modelling how
    declining chill hours under climate warming affect tree phenology. Probit and logistic usually give similar
    results; used in phenological-threshold studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 046_tobit_SSR_allele_freq_LOD_censored.xlsx
  >> SCENARIO (narration):
    When genotyping rare SSR alleles, low-frequency variants fall below the
    detection threshold so their LOD_under_censored values are left-censored at
    the limit of reliable scoring. We want to model the measured_frequency of
    each allele while accounting for the sample_count_N behind each estimate,
    recognizing that the censored LOD scores hide the true low values. Because
    the response is censored at a lower bound, Tobit Regression correctly
    handles the limit observations rather than discarding them. This yields
    unbiased estimates of how allele frequency relates to detection strength and
    sampling effort.
  >> VARIABLE SELECTION:
    - Outcome (censored): LOD_under_censored
    - Predictor: measured_frequency
    - Predictor: sample_count_N

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 2.27   Tobit (left-censored: frequency below detection limit LOD)
    Measured allele frequency ~ sample size

>> COMMENTARY (narration):
    A common problem in SSR (microsatellite) allele frequencies: very rare alleles fall below the limit of detection
    (LOD) -- flagged "not detected", but not zero. Deleting these or treating them as zero biases the result. Tobit
    regression correctly models this left-censored data. In genetic-diversity studies this is critical when estimating
    rare-allele frequencies: rare alleles are often near the LOD and carry the population's unique genetic heritage.
    Tobit is the statistically correct solution for working with "below detection limit" frequencies; it is superior
    to raw deletion/zeroing.

====================================================================================

#47  Bayesian Linear Regression
    file: 047_bayesian_linear_offspring_midparent_h2.xlsx
  >> SCENARIO (narration):
    In a half-sib heritability study we have only 22 families, and we want to
    estimate parent-offspring resemblance for stem size to infer narrow-sense
    heritability. We relate the half_sib_offspring_dbh to the midparent_dbh_cm,
    while noting each family's sample_count_family as a measure of confidence.
    With such a small sample, Bayesian Linear Regression lets us incorporate
    prior knowledge and produce full posterior uncertainty around the regression
    slope. The slope's posterior directly informs the heritability of diameter
    growth.
  >> VARIABLE SELECTION:
    - Outcome: half_sib_offspring_dbh
    - Predictor: midparent_dbh_cm
    - Predictor: sample_count_family

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Conjugate Normal-Inverse-Gamma prior   posterior sigma2 = 11.23
    Half-sib offspring dbh ~ midparent dbh

>> COMMENTARY (narration):
    We predicted offspring diameter from midparent diameter with Bayesian linear regression -- a classic heritability
    (h2) setup: the parent-offspring regression slope gives narrow-sense heritability. The Bayesian advantage: the
    result is not a single point but a probability distribution (posterior) and credible interval for the slope --
    honestly expressing heritability's uncertainty. The posterior error variance is sigma2 = 11.23. In forest genetics
    heritability determines the genetic gain achievable by selection; the Bayesian framework is a powerful way,
    especially in few-family studies, to show the uncertainty of an h2 estimate.

====================================================================================

#48  Nonlinear Regression
    file: 048_nonlinear_logistik_growth_curve_height_age.xlsx
  >> SCENARIO (narration):
    To characterize the growth trajectory of a forest stand, we measured tree
    height_m at successive age_year values, expecting growth to accelerate then
    plateau toward a maximum. We want to fit a logistic growth curve so the
    parameters carry biological meaning such as asymptotic height and the age of
    fastest growth. Because the relationship between height and age is
    inherently S-shaped and not linear, Nonlinear Regression with a logistic
    function is required. This produces a smooth growth model useful for yield
    prediction and rotation planning.
  >> VARIABLE SELECTION:
    - Outcome: height_m
    - Predictor: age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: y = K / (1 + exp(-r·(x-x0)))  (logistic growth)   R2 = 0.972
    Height ~ age

>> COMMENTARY (narration):
    We modelled how tree height changes with age. Tree growth is not linear: it rises fast when young, then saturates
    toward a ceiling (K -- the asymptotic maximum height) -- a classic S-curve (logistic growth). We modelled this
    with a logistic function and the fit is excellent (R2 = 0.97). The curve gives three parameters: the ceiling (K,
    the maximum height a species/provenance can reach), growth rate (r) and inflection point (x0). These are compared
    across provenances/clones to answer "which genotype grows faster/taller". Nonlinear regression is the basic tool
    of growth-yield modelling in forest mensuration; a straight line would misrepresent the nature of tree growth.

====================================================================================

#49  Ridge Regression
    file: 049_ridge_multicollinear_SNP_BV_wsg.xlsx
  >> SCENARIO (narration):
    In a genomic prediction effort we have nine SNP markers, SNP_1 through
    SNP_9, that are multicollinear because they sit in linkage disequilibrium
    across the same chromosomal region. We want to predict the wood specific
    gravity breeding value, BV_wsg_BLUP, from these correlated SNPs for marker-
    assisted selection. Because severe multicollinearity inflates ordinary
    regression coefficients, Ridge Regression stabilizes the estimates by
    shrinking them while keeping all markers in the model. This delivers a
    reliable predictive equation linking SNP genotypes to wood density breeding
    value.
  >> VARIABLE SELECTION:
    - Predictor: SNP_1
    - Predictor: SNP_2
    - Predictor: SNP_3
    - Predictor: SNP_4
    - Predictor: SNP_5
    - Predictor: SNP_6
    - Predictor: SNP_7
    - Predictor: SNP_8
    - Predictor: SNP_9
    - Outcome: BV_wsg_BLUP

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Ridge (alpha = 1.0): R2 = 0.890   Adj. R2 = 0.883   n = 150
    Wood-density breeding value (BV) ~ 9 SNPs (mutually correlated)

>> COMMENTARY (narration):
    Predicting the wood-density breeding value (BV BLUP) from 9 highly correlated SNP markers, ordinary regression
    becomes unstable (collinearity inflates coefficients -- SNPs are correlated due to linkage disequilibrium/LD).
    Ridge regression solves this by shrinking all coefficients -- it does not zero them but limits their magnitude,
    keeping prediction stable (R2 = 0.89). Ridge is ideal when "I want to keep all SNPs but there is collinearity". In
    genomic selection, predicting breeding values from thousands of correlated SNPs, Ridge (and its relative rrBLUP)
    are among the core methods.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 050_lasso_QTL_choice_3gercek_12gurultu.xlsx
  >> SCENARIO (narration):
    In a QTL-mapping experiment we suspect only three of many markers truly
    influence stem diameter, while the rest are noise. We have three real
    signals, QTL_1_actual, QTL_2_actual and QTL_3_actual, embedded among twelve
    noise markers, marker_noise_1 through marker_noise_12, and we want to
    predict phenotype_dbh_cm. Because we need automatic variable selection that
    drives irrelevant coefficients to exactly zero, Lasso Regression is the
    ideal tool. This identifies the truly associated QTLs and yields a
    parsimonious predictive model for diameter.
  >> VARIABLE SELECTION:
    - Predictor: QTL_1_actual
    - Predictor: QTL_2_actual
    - Predictor: QTL_3_actual
    - Predictor: marker_noise_1
    - Predictor: marker_noise_2
    - Predictor: marker_noise_3
    - Predictor: marker_noise_4
    - Predictor: marker_noise_5
    - Predictor: marker_noise_6
    - Predictor: marker_noise_7
    - Predictor: marker_noise_8
    - Predictor: marker_noise_9
    - Predictor: marker_noise_10
    - Predictor: marker_noise_11
    - Predictor: marker_noise_12
    - Outcome: phenotype_dbh_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Lasso (alpha = 0.1): R2 = 0.928   Adj. R2 = 0.920   n = 150
    12 of 15 markers shrunk to zero (automatic QTL selection)

>> COMMENTARY (narration):
    In this data 3 real QTLs (quantitative trait loci) were mixed with 12 noise markers. The strength of Lasso shows
    here: the penalty term drives the coefficients of irrelevant markers exactly to zero -- automatic QTL/marker
    selection. 12 of 15 markers were eliminated (leaving exactly the 3 real QTLs!), and the model is strong at R2 =
    0.93. Lasso is extremely useful when you want to pick the markers truly associated with the phenotype from many
    candidates -- in QTL mapping and genome-wide association (GWAS) studies. While Ridge shrinks all markers, Lasso
    eliminates the irrelevant ones entirely.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 051_mediation_temperature_bud_height_kaskad.xlsx
  >> SCENARIO (narration):
    In this provenance trial we ask whether spring temperature shapes seedling
    height growth directly, or indirectly through phenology. We hypothesize that
    warmer conditions advance bud burst, which in turn drives growth. To test
    this, we treat X_mean_temperature_C as the predictor, M_bud_fracture_jd (the
    Julian day of bud fracture) as the mediator, and Y_annual_height_growth_cm
    as the outcome. Mediation Analysis fits perfectly here because it decomposes
    the total temperature effect into a direct path and an indirect path running
    through bud phenology, telling us how much of the warming signal is
    channeled through timing of growth onset.
  >> VARIABLE SELECTION:
    - Predictor (X): X_mean_temperature_C
    - Mediator (M): M_bud_fracture_jd
    - Outcome (Y): Y_annual_height_growth_cm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 0.41   p < .001   95% CI (0.24 , 0.61)   Significant
    Path: temperature -> bud-flush day -> annual height growth

>> COMMENTARY (narration):
    We tested whether temperature's effect on height growth passes through bud-flush timing with mediation analysis.
    The indirect effect is significant (0.41, p < .001, CI excludes zero): temperature pulls the bud-flush day earlier,
    earlier flushing gives a longer growing season, which raises annual height growth. So temperature affects height
    not directly but "through" phenology (bud flush). This is quantified evidence of the climate-phenology-growth
    cascade -- the mechanism by which climate change affects tree growth. Mediation analysis opens up the intermediate
    mechanism in "why X affects Y".

====================================================================================

#52  Path Analysis
    file: 052_path_analysis_env_BV_survival_height.xlsx
  >> SCENARIO (narration):
    Here we want to understand how site environment and tree genetic quality
    jointly determine field performance in a forest planting. We propose that
    elevation_m and precipitation_mm influence both the breeding value BV_BLUP
    and downstream traits, while genetic merit and survival together shape five-
    year height. Using Path Analysis we can lay out a network in which
    elevation_m and precipitation_mm feed into BV_BLUP and
    survival_survival_ratio, which then propagate to height_cm_5yr. This
    approach is ideal because it estimates several interlinked regression paths
    at once, separating direct environmental effects from those mediated by
    genetics and survival.
  >> VARIABLE SELECTION:
    - Exogenous: elevation_m
    - Exogenous: precipitation_mm
    - Endogenous/Mediator: BV_BLUP
    - Endogenous/Mediator: survival_survival_ratio
    - Outcome: height_cm_5yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.721   RMSEA = 0.388   (poor fit -- model should be revised)
    Model: elevation/precipitation/BV -> height -> survival rate

>> COMMENTARY (narration):
    Path analysis is extended regression that models several cause-effect relationships at once. Here we tested that
    environmental variables and breeding value (BV) affect height, and that height determines survival. CFI = 0.72 and
    RMSEA = 0.39 -- the fit is poor, so the hypothesised model does not fully match the data; some paths may be missing
    or the relationships may not be linear. This is informative too: it says the theoretical causal model should be
    revised. Path analysis is a powerful way to separate direct and indirect effects and test environment-genetics-
    phenotype causal webs; both good and poor fit guide us in improving the model.

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 053_lmm_mixed_provenance_random_site_fixed.xlsx
  >> SCENARIO (narration):
    In this common-garden experiment we compare seedling height across multiple
    seed sources grown at several test sites. Because the six provenances were
    sampled from a broader pool while the four sites are the specific locations
    of interest, we treat them differently in the model. We fit a Linear Mixed
    Model with provenance_random as a random effect and site_fixed as a fixed
    effect, predicting height_cm_5yr. An LMM is the right tool because it
    partitions variance between provenances as a random source while still
    estimating fixed site contrasts on height.
  >> VARIABLE SELECTION:
    - Random effect: provenance_random
    - Fixed effect: site_fixed
    - Response: height_cm_5yr

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group variance (provenance random) = 0.0000   ICC = 0.0000
    Height ~ site (fixed)  +  (1 | provenance random)

>> COMMENTARY (narration):
    We analysed 5-year height with site as a fixed effect and provenance as a random effect in a mixed model. A
    striking result: the provenance random variance is zero (ICC = 0) -- in this model there is almost no
    between-provenance variance; the variability in height comes from within-provenance and the site effect. This is
    informative too: in this data/model combination provenance may add no extra variance once site is held fixed (or
    the provenance effect overlaps with the site effect). LMM is the correct framework for separating fixed and random
    effects; a zero variance component is a genuine result and signals the random effect may be unnecessary. It is the
    standard for modelling provenance/family random effects in forest trials.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 054_multiple_imputation_SNP_call_missing_MAR.xlsx
  >> SCENARIO (narration):
    Our genomic dataset has missing SNP dosage calls scattered across samples,
    and dropping incomplete trees would shrink our sample and bias diversity
    estimates. Assuming the gaps are missing at random, we want to recover
    usable values for SNP_1_dosage, SNP_2_dosage, and SNP_3_dosage so we can
    relate them to dbh_cm. Multiple Imputation is the appropriate method because
    it generates several plausible completed datasets, propagating the
    uncertainty of the missing genotype calls rather than relying on a single
    guess. This lets us analyze the marker-growth relationship without
    discarding partially genotyped individuals.
  >> VARIABLE SELECTION:
    - Variable with missing: SNP_1_dosage
    - Variable with missing: SNP_2_dosage
    - Variable with missing: SNP_3_dosage
    - Auxiliary/Outcome: dbh_cm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (number of imputations) = 5
    Missing SNP calls (MAR) estimated from dbh and other SNP dosages

>> COMMENTARY (narration):
    In genotyping data, missing SNP calls are inevitable (low quality, technical error); deleting them loses
    individuals/markers and biases results. Multiple imputation estimates the missing SNP dosages from other SNPs and
    the phenotype, creating 5 separate completed datasets, runs the analysis in each and combines the results --
    thereby also accounting for estimation uncertainty. Under the MAR (missing at random) assumption this is the
    modern missing-data standard. In genomic studies, missing-genotype imputation (especially reference-panel based)
    is a routine pre-processing step; this method shows its statistical basis.

====================================================================================

#55  GEE
    file: 055_gee_repeated_fidanlar_irrigation_annual.xlsx
  >> SCENARIO (narration):
    In this nursery trial we tracked the same seedlings over several years to
    test whether supplemental irrigation accelerates growth. Because each
    seedling contributes repeated height measurements across year, the
    observations within a plant are correlated and cannot be treated as
    independent. We use a GEE to model height_cm as a function of treatment
    (irrigated versus control) and year, accounting for the within-seedling
    correlation structure. GEE fits this longitudinal design well because it
    gives robust population-averaged estimates of the irrigation effect despite
    the repeated measures.
  >> VARIABLE SELECTION:
    - Time/Repeated: year
    - Predictor: treatment
    - Response: height_cm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    year coefficient = 18.23   p < .001 ***   QIC = 254.49
    Height ~ year  (repeated measures, clustered within seedling)

>> COMMENTARY (narration):
    We took repeated (annual) height measurements from the same seedlings -- one seedling's measurements are dependent.
    GEE estimates the relationship at the "population average" level in such clustered data, accounting for the
    correlation structure. The year coefficient is highly significant and large (18.23, p < .001): seedlings grow on
    average ~18 cm per year -- a strong, expected growth trend. GEE is an alternative to LMM: rather than modelling
    random effects, it corrects the correlation among observations and gives the marginal (average) effect. It is a
    common, powerful method for analysing repeated measures in longitudinal growth/irrigation trials.

====================================================================================

#56  GLMM
    file: 056_glmm_poisson_branch_count_provenance_site.xlsx
  >> SCENARIO (narration):
    We are studying branching architecture across seed sources planted at
    different sites, counting the side branches on each young tree. Since
    side_branch_count is a count outcome and trees are clustered within
    provenance and site, ordinary regression would mishandle both the
    distribution and the grouping. We fit a Poisson GLMM with provenance and
    site effects predicting side_branch_count. This model suits the data because
    it links the count response through a Poisson process while including random
    grouping so the provenance and site clustering is properly captured.
  >> VARIABLE SELECTION:
    - Grouping/Fixed: provenance
    - Grouping/Fixed: site
    - Count response: side_branch_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    site[S3] coefficient = -0.371, p < .001 ***
    Poisson GLMM: side-branch count ~ site  +  (1 | provenance)

>> COMMENTARY (narration):
    We modelled side-branch counts (Poisson) with both a site fixed effect and a provenance random effect -- a GLMM:
    mixed model + count distribution. One site's (S3) effect is significant and negative (coef -0.37, p < .001): at
    this site seedlings produce fewer side branches (environment affects branch development -- fewer branches usually
    means a straighter stem, positive for wood quality). GLMM combines the strengths of LMM (which assumes normality)
    and GLM (which ignores provenance clustering) for "repeated/nested count data". It is the correct analytic
    framework for forest-genetics data like branch/cone counts nested within provenances.

====================================================================================

#57  Elastic Net
    file: 057_elastic_net_SNP_2grup_correlated_wsg.xlsx
  >> SCENARIO (narration):
    We want to predict wood specific gravity from a large panel of SNP markers,
    but many of the genotype variables are strongly correlated within two marker
    blocks and several are pure noise. With so many collinear predictors, a
    single penalty would either drop or retain whole correlated groups
    arbitrarily. Elastic Net is ideal here because its blend of L1 and L2
    penalties handles the correlated SNP_group1 and SNP_group2 markers together
    while shrinking the SNP_noise variables, predicting wsg_gcm3 with a sparse,
    stable model.
  >> VARIABLE SELECTION:
    - Predictor: SNP_group1_1
    - Predictor: SNP_group1_2
    - Predictor: SNP_group1_3
    - Predictor: SNP_group1_4
    - Predictor: SNP_group2_1
    - Predictor: SNP_group2_2
    - Predictor: SNP_group2_3
    - Predictor: SNP_group2_4
    - Predictor: SNP_noise_1
    - Predictor: SNP_noise_2
    - Predictor: SNP_noise_3
    - Predictor: SNP_noise_4
    - Predictor: SNP_noise_5
    - Predictor: SNP_noise_6
    - Outcome: wsg_gcm3

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.001   R2 (train) = 0.902
    Wood density ~ 2 correlated SNP groups + noise

>> COMMENTARY (narration):
    Elastic Net is a hybrid of Ridge and Lasso: it both shrinks coefficients (like Ridge) and zeros out the irrelevant
    ones (like Lasso). Here there were two correlated SNP groups + noise; in such cases Lasso tends to pick one SNP
    from a group at random and drop the rest, whereas Elastic Net tends to keep the correlated group together. The
    model is strong (R2 = 0.90). In genomic selection, SNPs are correlated in blocks due to linkage disequilibrium
    (LD); Elastic Net evaluates these LD blocks together, providing both selection and stability. It is a balanced
    choice when there are correlated marker groups.

====================================================================================

#58  Robust Regression
    file: 058_robust_outlier_tree_measurement_height_age.xlsx
  >> SCENARIO (narration):
    We are modeling how tree height changes with age across a stand, but a
    handful of measurement errors and atypical trees produce extreme outliers in
    height_m. Ordinary least squares would let these aberrant points drag the
    fitted line and distort the age-height relationship. We therefore use Robust
    Regression to model height_m as a function of age_year, down-weighting the
    influential outliers. This method fits the situation because it yields a
    reliable growth slope that reflects the bulk of the trees rather than being
    skewed by a few erroneous records.
  >> VARIABLE SELECTION:
    - Predictor: age_year
    - Response: height_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust intercept = -0.033   vs   OLS intercept = 0.042   (outlier measurements affected OLS)
    Height ~ age

>> COMMENTARY (narration):
    This data had a few outlier tree measurements -- e.g. measurement error or unusual (forked/crushed) trees. Ordinary
    regression (OLS) takes outliers fully into account and is pulled toward them; the robust and OLS intercepts differ
    (-0.033 vs 0.042), showing the outliers' effect. Robust regression (Huber-T) gives less weight to outlying
    observations to preserve the true height-age relationship. In forest inventory, where outliers in field
    measurements are inevitable, robust regression gives a more reliable slope estimate than OLS.

====================================================================================

#59  Quantile Regression
    file: 059_quantile_heteroskedastik_height_age_provenance.xlsx
  >> SCENARIO (narration):
    In this dataset the spread of tree height widens with age and differs among
    seed sources, so the relationship is not captured well by a single average
    line. We are interested in how the upper and lower height percentiles, not
    just the mean, respond to age and provenance. Using Quantile Regression we
    model height_m against age_year and provenance across several quantiles.
    This approach fits the heteroskedastic structure because it describes how
    the tallest and shortest trees at each age diverge among provenances,
    revealing growth differences hidden in the conditional distribution.
  >> VARIABLE SELECTION:
    - Predictor: age_year
    - Predictor: provenance
    - Response: height_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    q = 0.10 slope = 0.750   |   q = 0.50 (median)   |   q = 0.90  (changing spread)
    Height ~ age (heteroskedastic)

>> COMMENTARY (narration):
    Ordinary regression models only the mean; but in the age-height relationship the spread grows as age increases
    (heteroskedasticity) -- young trees are similar in height, old trees very variable (some provenances much taller).
    Quantile regression models the lower (q=0.10, shortest trees), middle (q=0.50) and upper (q=0.90, tallest trees)
    quantiles separately. A different slope at the upper quantile means "the fastest-growing trees relate to age
    differently". In forest genetics, when the extremes (best/worst individuals) matter rather than the average -- as
    in plus-tree selection -- quantile regression reveals what mean-based models miss.

====================================================================================

#60  ROC Curve Analysis
    file: 060_roc_drought_tolerance_siniflandirici.xlsx
  >> SCENARIO (narration):
    We want to know how well physiological stress markers can classify seedlings
    as drought tolerant or sensitive. Each seedling has a proline_score, a
    chlorophyll_retention value, and an MDA_lipid_perox measure, with the known
    status recorded in drought_tolerant_1_0. ROC Curve Analysis is the right
    tool because it evaluates each marker as a continuous classifier against the
    binary outcome, plotting sensitivity versus specificity across thresholds.
    This lets us compare the markers' discriminating power and choose cutoffs
    for identifying drought-tolerant stock.
  >> VARIABLE SELECTION:
    - Classifier: proline_score
    - Classifier: chlorophyll_retention
    - Classifier: MDA_lipid_perox
    - Binary outcome: drought_tolerant_1_0

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.772 (acceptable discriminating power)   n = 230 (positive 65, negative 165)
    Drought-tolerance classifier (proline, chlorophyll, MDA)

>> COMMENTARY (narration):
    We measured the discriminating power of a classifier predicting whether a seedling is drought-tolerant from
    physiological biomarkers (proline, chlorophyll retention, MDA lipid peroxidation) with the ROC curve. AUC = 0.77
    is "acceptable": the biomarkers reasonably separate tolerant from sensitive seedlings (not perfect -- drought
    tolerance is multi-factorial, no single biomarker suffices). Because positives are rare (65/230), AUC is more
    informative than accuracy. ROC/AUC is ideal for evaluating the quality of biomarker-based screening models for
    drought-resistant genotype selection and climate-resilient seedling production.

====================================================================================

#61  True Skill Statistic (TSS)
    file: 061_tss_Salix_caprea_distribution_model.xlsx
  >> SCENARIO (narration):
    In this study we model the geographic distribution of goat willow, Salix
    caprea, across a mountainous landscape to support seed source mapping and
    assisted migration planning. For each location_id we recorded environmental
    predictors, elevation_m, precipitation_mm and temperature_C, together with
    the observed presence or absence of the species coded in
    Salix_caprea_present. We want to know how well our presence-absence model
    discriminates suitable from unsuitable habitat once threshold effects are
    accounted for. The True Skill Statistic is the right metric here because it
    summarizes sensitivity and specificity into a single threshold-independent
    score that is insensitive to species prevalence.
  >> VARIABLE SELECTION:
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Predictor: temperature_C
    - Observed: Salix_caprea_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.426 (good)   Sensitivity 0.60   Specificity 0.83   n = 200
    Salix caprea species distribution model (SDM)

>> COMMENTARY (narration):
    TSS (True Skill Statistic) is a common performance measure in species distribution models (SDM): sensitivity +
    specificity - 1. The model predicting Salix caprea (goat willow) presence from environmental variables (elevation,
    precipitation, temperature) has TSS = 0.43 -- "good": specificity high (0.83, separates absences well), sensitivity
    moderate (0.60). TSS's advantage is insensitivity to prevalence (how common the species is), so it is preferred in
    ecology as an alternative to AUC. For predicting how a species' range will shift under climate change and for
    setting conservation priorities, SDM and TSS are indispensable in forest ecology.

====================================================================================

#62  Confusion Matrix Metrics
    file: 062_complexity_matrix_drought_2sinif.xlsx
  >> SCENARIO (narration):
    Here we evaluate a classifier that flags drought-stressed seedlings so that
    vulnerable provenances can be screened in the nursery. For each seedling_id
    we have physiological measurements, proline_score and chlorophyll, alongside
    the model's continuous prediction_score, its discrete prediction_label_0_1,
    and the ground-truth actual_label_0_1. The goal is to quantify how often the
    model's drought calls are correct versus missed or falsely raised. Confusion
    Matrix Metrics are appropriate because they cross-tabulate predicted against
    actual labels to yield accuracy, precision, recall and related skill
    measures.
  >> VARIABLE SELECTION:
    - Predicted score: prediction_score
    - Predicted label: prediction_label_0_1
    - Actual label: actual_label_0_1
    - Feature: proline_score
    - Feature: chlorophyll

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.70   F1 = 0.511
    Drought 2-class prediction -- actual vs predicted label

>> COMMENTARY (narration):
    We broke down a drought-tolerance classifier's predictions with a confusion matrix. Although accuracy looks
    reasonable (70%), F1 = 0.51 is lower -- this means the precision/recall balance is moderate, and the model
    struggles with one class (probably the rarer tolerant class). The confusion matrix reveals what the "accuracy"
    number hides -- where each class is misclassified. The F1 score is more informative than accuracy especially with
    imbalanced classes. In forest genetics it is the indispensable tool for evaluating the real performance of
    physiology-based tolerance/resistance classification models.

====================================================================================

#63  Random Forest
    file: 063_random_forest_clone_identity_prediction.xlsx
  >> SCENARIO (narration):
    This analysis aims to fingerprint poplar clones from their growth and wood
    traits so that mislabeled ramets in a clonal trial can be identified. The
    categorical clone identity (clone_A through clone_D) is the response we want
    to predict from height_cm, wsg_gcm3, branch_count, dbh_cm and age_year.
    Because these predictors interact nonlinearly and clones overlap in any
    single trait, we need a flexible ensemble. A Random Forest fits naturally,
    combining many decision trees to classify clone identity while ranking which
    growth and wood-density variables drive the separation.
  >> VARIABLE SELECTION:
    - Target: clone
    - Predictor: height_cm
    - Predictor: wsg_gcm3
    - Predictor: branch_count
    - Predictor: dbh_cm
    - Predictor: age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.483   (clone identity prediction -- multi-class)
    Predicted from 5 phenotypic variables

>> COMMENTARY (narration):
    We tried to predict which clone a tree belongs to from five phenotypic variables (height, wood density, branch
    count, diameter, age) with Random Forest -- but accuracy is only 48%. This is instructive: clone identity is
    multi-class (many clones) and clones overlap heavily in phenotype -- i.e. it is hard to tell clones apart by
    morphology alone. This shows why molecular markers (DNA fingerprinting) are needed for clone identification;
    phenotype alone does not suffice. Random Forest's low accuracy reveals the difficulty of the problem (high
    phenotypic plasticity, between-clone similarity) -- itself a valuable finding.

====================================================================================

#64  SVM
    file: 064_svm_drought_tolerance_4ozellik.xlsx
  >> SCENARIO (narration):
    We want to separate drought-tolerant from drought-sensitive seedlings using
    a compact set of physiological markers for early provenance screening. Each
    seedling_id carries four features, proline, chlorophyll, MDA_perox and
    K_leaf_water, and is labeled tolerant or sensitive in label_tolerance. The
    question is whether a clean decision boundary exists in this four-
    dimensional physiological space. A Support Vector Machine suits this two-
    class problem because it finds the maximum-margin hyperplane separating
    tolerant and sensitive seedlings even when the classes are only modestly
    separated.
  >> VARIABLE SELECTION:
    - Target: label_tolerance
    - Predictor: proline
    - Predictor: chlorophyll
    - Predictor: MDA_perox
    - Predictor: K_leaf_water

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.818   (drought tolerance yes/no)
    Predicted from 4 physiological traits (proline, chlorophyll, MDA, leaf water)

>> COMMENTARY (narration):
    We found the best boundary separating drought-tolerant from sensitive seedlings with SVM. SVM seeks the maximum-
    margin separating surface between two classes and, via the kernel trick, can model nonlinear boundaries too.
    Accuracy is high at 81.8% -- four physiological traits (proline, chlorophyll, MDA, leaf water potential) separate
    tolerance status well. While ROC (#60) handles the same problem in terms of probability/threshold, SVM directly
    finds the best class boundary. SVM is especially powerful in small-to-medium, high-dimensional classification and
    is robust thanks to margin maximisation. It is a valuable tool in drought-resistant genotype screening.

====================================================================================

#65  Gradient Boosting
    file: 065_gradient_boosting_height_nonlinear_BVxEnv.xlsx
  >> SCENARIO (narration):
    This study predicts five-year height growth of poplar clones from genetics
    and environment to capture nonlinear genotype-by-environment interaction.
    The outcome height_cm_5yr is modeled from the categorical clone (clone_A to
    clone_D), the site variables elevation_m, precipitation_mm and
    temperature_C, and each genotype's parent_BV breeding value. We expect
    growth to respond to climate in a curved, interacting fashion rather than a
    simple linear way. Gradient Boosting is well suited because it builds
    additive trees that capture these nonlinear breeding-value-by-environment
    effects and improves prediction stage by stage.
  >> VARIABLE SELECTION:
    - Target: height_cm_5yr
    - Predictor: clone
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Predictor: temperature_C
    - Predictor: parent_BV

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.286   (height 3-tier class: low/medium/high)
    Height-tier prediction from environment + parent BV

>> COMMENTARY (narration):
    We split 5-year height into three tiers (low/medium/high) and predicted the tier from environmental variables and
    parent breeding value (BV) with Gradient Boosting. Accuracy is only 28.6% -- near chance level (33%) for a
    three-class problem, so the model weakly predicts the height tier. This is an instructive result: height is not
    explained by the SIMPLE sum of environment and parent BV; the strong Genotype×Environment interaction (seen in #7,
    #17, #109) and measurement noise make the tier hard to predict. Boosting is a powerful method that adds trees
    sequentially chasing the error, but if the signal is weak it cannot give high accuracy -- this itself reveals the
    complexity of G×E and the multi-factorial nature of height.

====================================================================================

#66  K-Means Clustering
    file: 066_kmeans_population_yapilanmasi_3grup_allele.xlsx
  >> SCENARIO (narration):
    Here we ask whether allele frequency profiles alone can recover the
    underlying population structure of our forest tree samples. For each
    individual we have five marker frequencies, freq_allele_A, freq_allele_B,
    freq_allele_C, freq_allele_D and freq_allele_E, and a known
    actual_population label (anatolia, Marmara, black_sea) held out for
    validation. The aim is to discover natural genetic groupings without using
    that label during clustering. K-Means is appropriate because it partitions
    individuals into a preset number of clusters by minimizing within-group
    variance in allele-frequency space, which we then compare against the true
    populations.
  >> VARIABLE SELECTION:
    - Feature: freq_allele_A
    - Feature: freq_allele_B
    - Feature: freq_allele_C
    - Feature: freq_allele_D
    - Feature: freq_allele_E
    - Validation: actual_population

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   Silhouette = 0.354 (moderate cluster structure)
    Population structure from 5 allele frequencies

>> COMMENTARY (narration):
    We split individuals into three groups by allele frequencies with K-Means -- a way to detect "population structure"
    in population genetics (similar to STRUCTURE/ADMIXTURE logic). Silhouette = 0.35 is moderate -- three genetic
    groups separate but not sharply, as expected in natural populations (gene flow blurs the groups). The three
    clusters indicate three genetic lineages/ancestries in the sample. K-Means finds natural groups in unlabeled
    genetic data; it is practical for identifying population structure, hybrid zones and conservation units (MU/ESU).
    A moderate silhouette may point to gene flow among groups.

====================================================================================

#67  Hierarchical Clustering
    file: 067_hierarchical_Quercus_filogenetik_UPGMA.xlsx
  >> SCENARIO (narration):
    This analysis reconstructs phylogenetic relationships among oak (Quercus)
    accessions by combining chloroplast DNA sequence variation with leaf
    morphology. For each type we scored eight cpDNA positions, cpDNA_pos_1
    through cpDNA_pos_8, alongside morphometric traits leaf_length_mm,
    leaf_width_mm, petiol_length_mm, lob_count, leaf_index and
    gland_density_mm2. We want to see which accessions are most similar and how
    they nest into a branching hierarchy. Hierarchical Clustering with UPGMA
    fits because it builds a dendrogram from pairwise distances, revealing the
    nested structure expected of a phylogeny.
  >> VARIABLE SELECTION:
    - Feature: cpDNA_pos_1_A_C_G_T
    - Feature: cpDNA_pos_2_A_C_G_T
    - Feature: cpDNA_pos_3_A_C_G_T
    - Feature: cpDNA_pos_4_A_C_G_T
    - Feature: cpDNA_pos_5_A_C_G_T
    - Feature: cpDNA_pos_6_A_C_G_T
    - Feature: cpDNA_pos_7_A_C_G_T
    - Feature: cpDNA_pos_8_A_C_G_T
    - Feature: leaf_length_mm
    - Feature: leaf_width_mm
    - Feature: petiol_length_mm
    - Feature: lob_count
    - Feature: leaf_index
    - Feature: gland_density_mm2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 6   Silhouette = 0.164 (weak/overlapping structure)
    Quercus cpDNA + morphology (UPGMA phylogeny)

>> COMMENTARY (narration):
    We grouped Quercus (oak) species by cpDNA sequences and morphological traits with hierarchical clustering (UPGMA)
    -- the result is a dendrogram (phylogenetic tree). Splitting into six clusters gives a silhouette of only 0.16:
    the clusters overlap, there is no sharp separation. This is very typical in oaks -- they have a blurry phylogeny
    due to widespread hybridisation and incomplete lineage sorting; species do not separate cleanly genetically. The
    low silhouette reflects this biological reality (weak grouping = strong hybridisation/gene flow). Hierarchical
    clustering is the basic tool for seeing phylogenetic relationships and inter-species permeability.

====================================================================================

#68  DBSCAN Clustering
    file: 068_dbscan_population_cluster_koord_allele.xlsx
  >> SCENARIO (narration):
    We want to detect spatial genetic clusters of sampled trees while leaving
    isolated outliers unassigned, to delineate distinct populations across the
    landscape. For each population_id we have geographic coordinates lat and lon
    plus the frequency of the focal marker freq_allele_A. The goal is to find
    dense aggregations of genetically similar, geographically close individuals
    without fixing the number of groups in advance. DBSCAN is the right tool
    because it groups points by density and flags sparse points as noise,
    naturally separating core populations from scattered individuals.
  >> VARIABLE SELECTION:
    - Feature: lat
    - Feature: lon
    - Feature: freq_allele_A

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   Silhouette = 0.864   (population locations)
    Populations clustered by geographic density

>> COMMENTARY (narration):
    We clustered populations' geographic locations by density with DBSCAN. Unlike K-Means: we do not give the number
    of clusters in advance, and it flags low-density isolated populations as "noise" (outliers). Three dense population
    groups were found (silhouette 0.86, very strong). This shows populations cluster in particular geographic regions.
    In conservation genetics this is highly valuable: geographically isolated/sparse populations (noise points)
    usually have unique genetic heritage and should be prioritised for conservation. DBSCAN is more appropriate than
    K-Means for spatial population data with irregularly shaped clusters and isolated points.

====================================================================================

#69  PCA
    file: 069_pca_PROC_PRINCOMP_10SSR_He.xlsx
  >> SCENARIO (narration):
    This study summarizes the genetic diversity captured by ten SSR markers
    across five tree populations to visualize how they differentiate. For each
    subject we have the population label (Saros_gulf, Goksu_delta,
    kizilirmak_Avanos, sakarya_Adapazari, Goksun_valley) and expected
    heterozygosity values SSR_01_He through SSR_10_He. We want to compress these
    ten correlated marker dimensions into a few axes that explain most of the
    variation among individuals. Principal Component Analysis fits because it
    produces orthogonal components ordered by variance, letting us project
    populations into a low-dimensional diversity space.
  >> VARIABLE SELECTION:
    - Feature: SSR_01_He
    - Feature: SSR_02_He
    - Feature: SSR_03_He
    - Feature: SSR_04_He
    - Feature: SSR_05_He
    - Feature: SSR_06_He
    - Feature: SSR_07_He
    - Feature: SSR_08_He
    - Feature: SSR_09_He
    - Feature: SSR_10_He
    - Grouping: population

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 = 50.0%   PC2 = 41.6%   (first two components total 91.6%)
    10 SSR loci heterozygosity (He)

>> COMMENTARY (narration):
    We reduced the heterozygosity of 10 SSR (microsatellite) loci to a few summary axes with PCA. The first two
    components explain 91.6% of the total variance (PC1 50%, PC2 41.6%) -- we can summarise ten loci of genetic-
    diversity information in almost two axes. In population genetics PCA is the most common way to visualise
    individuals/populations in a two-dimensional space by genetic similarity -- to see population structure,
    hybridisation and gene flow (a fast, assumption-free alternative to STRUCTURE). The two components being this
    dominant shows the genetic diversity is gathered into a strong, low-dimensional structure.

====================================================================================

#70  t-SNE
    file: 070_tsne_filogenetik_haplotype_16SNP.xlsx
  >> SCENARIO (narration):
    Here we visualize haplotype structure among individuals genotyped at sixteen
    chloroplast SNPs to reveal phylogeographic groupings. Each subject carries a
    population_group label (europe, anatolia, iran) and SNP values cpDNA_SNP_01
    through cpDNA_SNP_16. The aim is to embed this high-dimensional SNP profile
    into two dimensions while preserving local similarity between haplotypes.
    t-SNE is appropriate because it is designed to map high-dimensional points
    into a low-dimensional layout that keeps nearby genotypes close, exposing
    clusters that align with phylogeographic origin.
  >> VARIABLE SELECTION:
    - Feature: cpDNA_SNP_01
    - Feature: cpDNA_SNP_02
    - Feature: cpDNA_SNP_03
    - Feature: cpDNA_SNP_04
    - Feature: cpDNA_SNP_05
    - Feature: cpDNA_SNP_06
    - Feature: cpDNA_SNP_07
    - Feature: cpDNA_SNP_08
    - Feature: cpDNA_SNP_09
    - Feature: cpDNA_SNP_10
    - Feature: cpDNA_SNP_11
    - Feature: cpDNA_SNP_12
    - Feature: cpDNA_SNP_13
    - Feature: cpDNA_SNP_14
    - Feature: cpDNA_SNP_15
    - Feature: cpDNA_SNP_16
    - Grouping: population_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence (fit quality) = 0.177 -> very good
    16 cpDNA SNP haplotypes reduced to 2D

>> COMMENTARY (narration):
    t-SNE is a nonlinear method that projects high-dimensional haplotype data (16 cpDNA SNPs) into a two-dimensional
    plot. Unlike PCA, it focuses on preserving local neighbourhoods -- similar haplotypes are placed close on the map
    -- ideal for seeing hidden phylogenetic/lineage structure. The KL divergence is very low (0.18), so the reduction
    quality is very good; haplotype groups form clear clusters on the map. Important caveat: t-SNE axes and
    between-cluster distances are not interpretable, only the clustering pattern is. In phylogeography and haplotype-
    network analysis it is a powerful tool for visually exploring similar lineage groups.

====================================================================================

#71  MDS
    file: 071_mds_genetic_distance_population.xlsx
  >> SCENARIO (narration):
    We sampled 30 forest tree populations and characterized each with a suite of
    standard population-genetic indices: expected heterozygosity He, observed
    heterozygosity Ho, the fixation index Fst, the inbreeding coefficient Fis,
    allelic richness AR, and the count of private alleles PA. Our goal is to
    visualize how genetically similar or distant these populations are from one
    another based on their combined multivariate genetic profile. Because we
    want to represent the pairwise genetic distances among populations in a low-
    dimensional map that preserves those distances as faithfully as possible,
    Multidimensional Scaling (MDS) is the appropriate technique. The resulting
    ordination lets us see clusters of genetically alike populations and
    identify outliers.
  >> VARIABLE SELECTION:
    - Grouping: population
    - Indicator: He
    - Indicator: Ho
    - Indicator: Fst
    - Indicator: Fis
    - Indicator: AR
    - Indicator: PA

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.271 (very weak fit)
    Genetic distance among populations (He, Ho, Fst, Fis, AR, PA)

>> COMMENTARY (narration):
    MDS maps the genetic distance among populations (from diversity parameters -- He, Ho, Fst, Fis, allele richness,
    private alleles) into two dimensions -- populations with similar genetic profiles placed near, dissimilar ones far.
    Stress = 0.27 is a "very weak" fit; this means the population genetic structure does not fully fit two dimensions,
    it is more complex (multidimensional) -- common in population genetics. Still, the general gradient is visible. MDS
    is the classic way in population genetics to visualise genetic-distance (Nei, Fst) matrices; the stress value
    honestly tells how much we can trust the map.

====================================================================================

#72  UMAP
    file: 072_umap_population_marginal_center_gradient.xlsx
  >> SCENARIO (narration):
    We genotyped 225 trees drawn from three position classes along a species
    distribution gradient: marginal_east, center, and marginal_west populations,
    recording allele frequencies at three loci (freq_locus_1, freq_locus_2,
    freq_locus_3) plus the mean expected heterozygosity He_mean. We want to
    explore whether trees from marginal versus central populations separate
    naturally in genetic space, which would signal differentiation at the range
    edges. Since the structure may be highly non-linear and we want a faithful
    2D embedding that preserves local neighborhoods of similar genotypes, UMAP
    is well suited to reveal this clustering. Coloring the embedding by
    population_type shows whether the marginal and center groups occupy distinct
    regions.
  >> VARIABLE SELECTION:
    - Color group: population_type
    - Feature: freq_locus_1
    - Feature: freq_locus_2
    - Feature: freq_locus_3
    - Feature: He_mean

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    UMAP dimension reduction (locus frequencies + He -> 2D)
    n_neighbors and min_dist balance local/global structure

>> COMMENTARY (narration):
    UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure
    better and is faster. We projected population profiles of locus frequencies and heterozygosity into a two-
    dimensional map; similar genetic profiles cluster, and the marginal-central gradient can be visualised.
    n_neighbors tunes the local-global balance, min_dist the tightness of clusters. UMAP is increasingly preferred for
    discovering hidden population structure and gradients in high-dimensional genetic/genomic data. It is used for
    visualisation; the clustering pattern is interpreted, not the axis values.

====================================================================================

#73  Cronbach's Alpha
    file: 073_cronbach_alpha_tree_form_quality_8madde.xlsx
  >> SCENARIO (narration):
    To assess tree form quality, 160 raters scored individual trees on eight
    phenotypic items: form_straight, branch_angle_suitable, steep_trunk,
    one_leader, side_branch_silvik, large_age, healthy_bark, and leaflet_dense.
    We want to know whether these eight items consistently measure a single
    underlying construct of tree form quality before we combine them into a
    composite score. Cronbach's Alpha quantifies the internal consistency of
    this multi-item scale, telling us how reliably the items hang together. A
    high alpha would justify treating the eight items as one coherent quality
    index.
  >> VARIABLE SELECTION:
    - Item: form_straight
    - Item: branch_angle_suitable
    - Item: steep_trunk
    - Item: one_leader
    - Item: side_branch_silvik
    - Item: large_age
    - Item: healthy_bark
    - Item: leaflet_dense

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.945 (excellent internal consistency)   8 items
    Tree-form quality assessment scale

>> COMMENTARY (narration):
    We tested the internal consistency of an 8-item scale assessing tree-form quality (stem straightness, branch
    angle, single leader, healthy bark, etc.) with Cronbach's alpha. Alpha = 0.95 is excellent: the eight items
    consistently measure the same underlying concept (tree-form quality), and raters score coherently across items.
    Showing the reliability of a form-quality scoring scale before using it is essential -- low alpha signals the items
    do not measure the same thing and the total form-score may be meaningless. Above 0.90 is excellent. It is the
    first step of proving the reliability of a scoring scale in plus-tree selection and form-quality assessment.

====================================================================================

#74  Likert Analysis
    file: 074_likert_breeding_oncelik_5li_10madde.xlsx
  >> SCENARIO (narration):
    In a survey of 200 participants we asked them to rate ten breeding-priority
    statements (item_breeding_1 through item_breeding_10) on a five-point
    agreement scale, also recording each respondent's institution (TARGEM,
    university, ogm, private) and age_group. Our aim is to summarize how
    strongly stakeholders endorse each breeding priority and to display the
    distribution of agreement across the ten items. Likert analysis is the
    natural choice because it summarizes ordered-category responses item by
    item, showing the spread of agree/disagree levels. We can then compare
    response patterns by institution and age_group.
  >> VARIABLE SELECTION:
    - Likert item: item_breeding_1
    - Likert item: item_breeding_2
    - Likert item: item_breeding_3
    - Likert item: item_breeding_4
    - Likert item: item_breeding_5
    - Likert item: item_breeding_6
    - Likert item: item_breeding_7
    - Likert item: item_breeding_8
    - Likert item: item_breeding_9
    - Likert item: item_breeding_10
    - Group: institution
    - Group: age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.172 (UNACCEPTABLE)   10 items
    Breeding-priority 5-point Likert scale

>> COMMENTARY (narration):
    A 10-item, 5-point Likert scale measuring breeding priorities gave a striking result: Cronbach's alpha is only
    0.17 -- unacceptably low. This means the items do NOT measure a common concept; it is as if 10 unrelated
    priorities were asked. Actually this is expectable -- breeding goals (growth, quality, resistance) are inherently
    different, even competing priorities; summing them into a single "priority" dimension would be wrong. This is
    instructive: a group of items does not always form a single scale. A low alpha says the scale is not
    unidimensional and should be split into sub-dimensions or the items evaluated separately. Likert analysis exposes
    this diagnosis early.

====================================================================================

#75  EFA
    file: 075_efa_PROC_FACTOR_3faktor_12madde.xlsx
  >> SCENARIO (narration):
    We measured 180 trees on twelve traits spanning growth (growth_item_1 to
    growth_item_4), adaptation (adaptation_item_5 to adaptation_item_8), and
    wood quality (wood_quality_item_9 to wood_quality_item_12). We suspect these
    twelve observed variables are driven by a smaller number of latent
    dimensions, and we want to uncover how many factors underlie them and which
    items load on each. Exploratory Factor Analysis (EFA) is appropriate because
    we are letting the data reveal the latent factor structure rather than
    imposing one in advance. This should clarify whether growth, adaptation, and
    wood-quality traits form three distinct factors.
  >> VARIABLE SELECTION:
    - Item: growth_item_1
    - Item: growth_item_2
    - Item: growth_item_3
    - Item: growth_item_4
    - Item: adaptation_item_5
    - Item: adaptation_item_6
    - Item: adaptation_item_7
    - Item: adaptation_item_8
    - Item: wood_quality_item_9
    - Item: wood_quality_item_10
    - Item: wood_quality_item_11
    - Item: wood_quality_item_12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO (overall) = 0.830 (Very good)   -> factor analysis appropriate
    12 items, expected 3 factors (growth / adaptation / wood quality)

>> COMMENTARY (narration):
    We explored how many hidden dimensions (factors) underlie 12 tree-trait items with Exploratory Factor Analysis.
    KMO = 0.83 is "very good": the data are suitable for factor analysis, with enough shared variance among items. The
    items group into three factors as expected -- growth, adaptation, wood quality. EFA reduces many traits to a few
    meaningful dimensions, simplifying assessment and revealing the underlying structures. This is very valuable in
    tree breeding when building selection indices: separating traits into independent dimensions (growth vs quality vs
    resistance) is the basis of a multi-trait selection strategy.

====================================================================================

#76  ICC
    file: 076_icc_3lab_dbh_measurement_reliability.xlsx
  >> SCENARIO (narration):
    Three laboratories independently measured diameter at breast height on the
    same 30 trees, producing lab1_dbh_cm, lab2_dbh_cm, and lab3_dbh_cm. Before
    pooling DBH data across labs, we need to confirm that the three labs agree
    closely on the same tree. The Intraclass Correlation Coefficient (ICC)
    quantifies this measurement reliability by partitioning between-tree
    variance from between-lab error. A high ICC would assure us that DBH
    measurements are interchangeable across the three labs.
  >> VARIABLE SELECTION:
    - Rater measurement: lab1_dbh_cm
    - Rater measurement: lab2_dbh_cm
    - Rater measurement: lab3_dbh_cm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.994   (Excellent)   3 labs/measurers   95% CI ~ (0.99 , 1.00)

>> COMMENTARY (narration):
    We assessed how consistent three labs/measurers are in measuring the stem diameter (dbh) of the same trees with
    ICC. ICC = 0.994 is "excellent": inter-measurer agreement is almost perfect -- whichever measurer measures, the
    result is the same. This is critical in forest inventory and genetic trials: if diameter measurement varied from
    measurer to measurer, all growth data and genetic analyses would be suspect. ICC is the standard for measuring
    inter-measurer/lab reliability on continuous measurements (diameter, height) (Cohen's Kappa is for categorical
    data, ICC for continuous). It is the standard way to prove dendrometric measurement reliability.

====================================================================================

#77  CFA
    file: 077_cfa_calis_SEM_growth_stress_2faktor.xlsx
  >> SCENARIO (narration):
    Based on theory, we hypothesize that eight measured indicators in 200 trees
    reflect two latent constructs: a growth factor (growth_1 to growth_4) and a
    stress-response factor (stress_1 to stress_4), with trees drawn from six
    provenances and varying in age_year. Unlike exploratory work, here we
    already have a defined two-factor model and want to test how well it fits
    the observed data. Confirmatory Factor Analysis (CFA) within an SEM
    framework is the right tool because it evaluates a pre-specified measurement
    structure rather than discovering one. Good fit would validate growth and
    stress as separable latent dimensions.
  >> VARIABLE SELECTION:
    - Growth indicator: growth_1
    - Growth indicator: growth_2
    - Growth indicator: growth_3
    - Growth indicator: growth_4
    - Stress indicator: stress_1
    - Stress indicator: stress_2
    - Stress indicator: stress_3
    - Stress indicator: stress_4
    - Grouping: provenance
    - Covariate: age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.998   RMSEA = 0.023   (excellent fit)
    2 factors: growth + stress

>> COMMENTARY (narration):
    While EFA explores hidden structure, CFA tests a pre-specified theory: here we tested whether two distinct trait
    dimensions -- "growth" and "stress resistance" -- fit the data. The fit is excellent (CFI = 0.998, RMSEA = 0.023):
    the items really load onto these two dimensions, the hypothesised trait structure is fully confirmed. CFA is part
    of structural equation modelling (SEM) and is the strongest way to prove scale/structure validity. In tree
    breeding, confirming the structure of multi-dimensional trait complexes like growth and stress-resistance matters
    for the validity of selection indices and multi-trait breeding models.

====================================================================================

#78  Survey Means
    file: 078_survey_means_regional_He_population.xlsx
  >> SCENARIO (narration):
    In a complex survey of 200 populations stratified by region (anatolia,
    Marmara, mediterranean, black_sea) and organized into clusters with sampling
    weights, we recorded mean expected heterozygosity He_mean and allelic
    richness AR_allele_richness. We want to estimate the regional average
    genetic diversity while properly accounting for the stratified, clustered,
    weighted design. Survey Means analysis is required because ignoring weight,
    stratum_region, and cluster_id would bias the estimates and their standard
    errors. This yields design-consistent estimates of mean He and allele
    richness per region.
  >> VARIABLE SELECTION:
    - Stratum: stratum_region
    - Cluster: cluster_id
    - Weight: weight
    - Mean variable: He_mean
    - Mean variable: AR_allele_richness

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    He weighted mean = 0.327   SE = 0.020   95% CI (0.288 , 0.366)   CV 5.96%
    Stratified + clustered + weighted design (Taylor SE)

>> COMMENTARY (narration):
    In a regional genetic inventory, populations are not chosen at random but with a stratified, clustered and
    weighted design. A simple mean would be wrong for such data; complex-sample analysis gives the correct mean
    heterozygosity (He) and standard error (Taylor linearization) by accounting for the design weights and clustering.
    The regional He weighted mean is 0.327 (CV 5.96%). National forest-genetic-resource inventories (genetic-diversity
    monitoring) are always designed this way; survey methods are essential for population-generalisable, unbiased
    genetic-diversity estimation.

====================================================================================

#79  Survey Frequencies
    file: 079_survey_frequency_haplotype_distribution.xlsx
  >> SCENARIO (narration):
    Across 220 sampled populations in a stratified cluster survey (regions
    anatolia, mediterranean, black_sea, Marmara) with sampling weights, we
    recorded the dominant chloroplast haplotype of each population (H1 through
    H5). We want to estimate the population-level proportion of each haplotype
    while respecting the complex sampling design. Survey Frequencies analysis is
    appropriate because it produces weighted, design-corrected estimates of
    categorical proportions and their confidence intervals. This tells us how
    haplotype distribution varies across the surveyed regions.
  >> VARIABLE SELECTION:
    - Stratum: stratum_region
    - Cluster: cluster_id
    - Weight: weight
    - Category: haplotype

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H1 proportion = 0.292 (SE 0.071)   |   H2 proportion = 0.287 (SE 0.070)
    Haplotype distribution -- stratified + clustered + weighted design

>> COMMENTARY (narration):
    We estimated haplotype (cpDNA lineage marker) frequencies among populations, accounting for the complex sample
    design. The H1 and H2 haplotypes are at similar frequency (~29%), with design-based standard errors. A simple
    percentage would give the wrong standard error by ignoring stratification and clustering; survey frequency
    analysis produces a confidence interval reflecting the true uncertainty. Haplotype frequencies are critical for
    phylogeography and gene conservation (which lineage is common where). It is the correct way to estimate haplotype/
    genotype proportions in a population in a design-consistent manner.

====================================================================================

#80  Survey Totals
    file: 080_survey_total_active_population_Ne_total.xlsx
  >> SCENARIO (narration):
    In a stratified cluster survey of 180 populations across three strata (I,
    II, III) with sampling weights, we measured the effective population size,
    active_population_Ne, of each population. Our objective is to estimate the
    total effective population size aggregated across all surveyed populations
    while honoring the weighted, clustered design. Survey Totals analysis is the
    correct method because it produces a design-consistent population total with
    valid standard errors from the weights, strata, and clusters. The result
    gives us the overall Ne of the active breeding population.
  >> VARIABLE SELECTION:
    - Stratum: stratum
    - Cluster: cluster_id
    - Weight: weight
    - Total variable: active_population_Ne

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total effective population (Ne) = 91,756   SE = 5,672   95% CI (80,563 , 102,948)
    Stratified design (Taylor SE)

>> COMMENTARY (narration):
    We estimated the total effective population size (Ne -- the number of genetically breeding individuals) in a
    region using the sampling design and weights: about 91,756 (95% CI 80,563 - 102,948). Ne is the most important
    parameter in conservation genetics: low Ne means risk of genetic drift and inbreeding. Scaling from the sampled
    populations up to the whole region requires correct weighting. Survey total analysis does exactly this and
    expresses uncertainty with the standard error. It is the official method used to assess the long-term
    sustainability of genetic resources.

====================================================================================

#81  Survey Regression
    file: 081_survey_regression_He_elevation_weighted.xlsx
  >> SCENARIO (narration):
    In this study we examine how environmental gradients shape genetic diversity
    across oak populations sampled under a stratified design. Using a survey-
    weighted regression, we model He_mean as a function of elevation_m and
    precipitation_mm, while respecting the survey weight and the stratification
    by stratum and clustering by cluster_id. Because populations were drawn with
    unequal selection probabilities across strata, ignoring the weight would
    bias our estimates of how diversity responds to climate. This approach lets
    us draw population-representative conclusions about whether higher-
    elevation, wetter sites harbour greater expected heterozygosity.
  >> VARIABLE SELECTION:
    - Weight: weight
    - Stratum: stratum
    - Cluster: cluster_id
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Outcome: He_mean

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.211   elevation coefficient = -0.0001   p < .001 ***
    He ~ elevation + precipitation   (stratified + clustered + weighted)

>> COMMENTARY (narration):
    We modelled heterozygosity (He) from elevation and precipitation, accounting for the complex sample design (strata,
    clusters, weights). The elevation effect is significant and negative (coefficient -0.0001, p < .001): genetic
    diversity falls as altitude rises -- high-altitude populations (usually small, isolated, marginal) are less
    diverse. This is an important pattern in conservation genetics. The model explains 21% of the variance. Design-
    based standard errors are more accurate than the (often too optimistic) errors of simple OLS. When estimating
    relationships from national genetic-inventory data, survey regression is necessary for valid inference.

====================================================================================

#82  Survey Logistic Regression
    file: 082_survey_logistic_endemic_asset_elevation.xlsx
  >> SCENARIO (narration):
    Here we ask whether the presence of endemic forest taxa is driven by
    elevation across a stratified regional sample. We fit a survey logistic
    regression with endemic_present_1_0 as the binary outcome and elevation_m as
    the predictor, incorporating the survey weight, the stratum design variable
    and cluster_id to account for the complex sampling scheme. Since occurrence
    data were collected under disproportionate stratum sampling, weighting is
    essential to estimate the true elevational probability of endemism. The
    result tells us how the odds of finding an endemic species shift as we climb
    in altitude.
  >> VARIABLE SELECTION:
    - Weight: weight
    - Stratum: stratum
    - Cluster: cluster_id
    - Predictor: elevation_m
    - Outcome: endemic_present_1_0

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    elevation: OR = 1.0016   p < .001 ***
    Endemic-species presence ~ elevation   (stratified + clustered + weighted)

>> COMMENTARY (narration):
    Predicting the probability of an endemic species being present at a locality from elevation, we handled both the
    binary outcome (logistic) and the complex sample design together. The odds ratio for elevation is 1.0016
    (p < .001): the odds of endemic presence rise with altitude -- high-mountain habitats are endemism hot spots
    (isolation and special conditions favour endemics). Because the design weights and clustering are accounted for,
    the OR and confidence interval are design-consistent and valid. Survey logistic regression is the correct way to
    model binary outcomes (endemic present/absent) from complex biogeographic surveys; it is valuable for identifying
    biodiversity hot spots.

====================================================================================

#83  GAM
    file: 083_gam_smooth_age_elevation_height.xlsx
  >> SCENARIO (narration):
    We investigate how tree height develops as a smooth, potentially nonlinear
    function of age and site elevation. Using a generalized additive model, we
    let height_m respond flexibly to smooth terms for age_year and elevation_m
    rather than forcing a straight-line relationship. This matters because
    forest growth typically accelerates and then plateaus with age, and
    elevation can impose nonlinear environmental limits. The GAM reveals the
    shape of these growth curves without us pre-specifying their form.
  >> VARIABLE SELECTION:
    - Smooth predictor: age_year
    - Smooth predictor: elevation_m
    - Outcome: height_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (explained) = 0.956
    Height ~ s(age) + s(elevation)  (smooth functions)

>> COMMENTARY (narration):
    GAM is a flexible generalisation of linear regression: it models a variable's effect with "smooth curves"
    (splines) rather than a straight line. Height's relationship with age and elevation is nonlinear -- the age effect
    saturates (growth curve), the elevation effect may peak at an optimal altitude. GAM learns this curved pattern from
    the data without assuming its shape in advance; the model explains 96% of the variance (very high). It preserves
    interpretability while adding flexibility: we can read from the curves at which age height accelerates and at which
    elevation it is best. For nonlinear but interpretable growth-environment relationships in forest genetics, it is
    an ideal method.

====================================================================================

#84  Discriminant Analysis (LDA/QDA)
    file: 084_diskriminant_PROC_DISCRIM_Quercus_3tur.xlsx
  >> SCENARIO (narration):
    In this analysis we test whether three oak species can be reliably separated
    on the basis of leaf morphology. We apply discriminant analysis with type
    (Q_robur, Q_petraea, Q_cerris) as the grouping variable and leaf_length_mm,
    leaf_width_mm, petiol_ratio and lob_count as the discriminating
    measurements. Because leaf traits often overlap between closely related
    oaks, a multivariate classifier helps us find the combination of traits that
    best distinguishes the taxa. The output quantifies how accurately morphology
    assigns each tree to its true species.
  >> VARIABLE SELECTION:
    - Group: type
    - Predictor: leaf_length_mm
    - Predictor: leaf_width_mm
    - Predictor: petiol_ratio
    - Predictor: lob_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.904   (3 Quercus species)
    Separation from 4 leaf traits (leaf length, width, petiole ratio, lobe count)

>> COMMENTARY (narration):
    Discriminant analysis finds the linear combinations of variables that best separate the groups (three oak species)
    -- a bit like a classification-focused PCA. Accuracy is 90.4%: four leaf-morphology traits separate the three
    Quercus species with high accuracy (a good result in a hybridising genus like oak). Discriminant analysis both
    classifies and answers "which morphological trait matters most for the separation" (discriminant function
    loadings). In morphometric taxonomy -- distinguishing species by leaf measurements -- it is the classic choice. It
    is used to assign new samples to existing species groups and to detect hybrid individuals.

====================================================================================

#85  Conditional Logit
    file: 085_conditional_logit_seedling_choice_grower.xlsx
  >> SCENARIO (narration):
    This study explores how growers choose among competing seedling options
    offered side by side. Using a conditional logit model, we model the binary
    chosen indicator across alternatives nested within each grower_id, with
    growth_ratio, wsg_score_1_5 and price_TL_seedling as attributes of each
    seedling alternative. Conditional logit is the right tool because it
    compares attributes within each choice set rather than across individuals.
    The estimates reveal how growth performance, wood quality and price trade
    off in seedling purchasing decisions.
  >> VARIABLE SELECTION:
    - Choice set: grower_id
    - Alternative: alternative
    - Attribute: growth_ratio
    - Attribute: wsg_score_1_5
    - Attribute: price_TL_seedling
    - Chosen: chosen

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (McFadden) = 0.0014  (very low)
    Seedling choice ~ growth rate + wood score + seedling price

>> COMMENTARY (narration):
    We examined which seedling growers choose with a conditional logit model -- each grower selects one from a choice
    set. McFadden pseudo R2 = 0.0014 is very low -- an instructive result: the chosen variables (growth rate, wood
    score, price) barely explain growers' choices. So the choice is driven not by these measured factors but by other
    (unmeasured) ones -- tradition, brand, availability, personal preference. A low pseudo R2 is also a valuable
    finding: it says "the variables in the model are insufficient". Conditional logit is the basis of discrete-choice
    analysis; the low fit shows the choice behaviour is more complex.

====================================================================================

#86  Kaplan-Meier
    file: 086_kaplan_meier_seedling_irrigation_survival_survival.xlsx
  >> SCENARIO (narration):
    We assess whether supplemental irrigation improves seedling survival under
    field establishment. With a Kaplan-Meier analysis we estimate survival over
    duration_month, using death_event to mark mortality and comparing the
    irrigated versus control treatment groups, while also exploring differences
    across provenance origins. This nonparametric approach handles censored
    seedlings that were still alive at the end of monitoring without assuming
    any particular survival distribution. The resulting survival curves show how
    long seedlings persist under each watering regime.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death_event
    - Group: treatment
    - Group: provenance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 16.5 months   (grouped by irrigation treatment)
    Log-rank test compares groups

>> COMMENTARY (narration):
    We analysed seedlings' survival time -- by irrigation treatment -- with Kaplan-Meier. Median survival is 16.5
    months. Kaplan-Meier analyses "time-to-event" data -- correctly using not-yet-dead (censored) seedlings too -- and
    tests treatment-group differences in survival with the log-rank test. In forest genetics, seedling survival is the
    basic indicator of afforestation success and of provenance/treatment suitability. It is the basic visual and test
    for evaluating the effect of interventions like irrigation on seedling survival.

====================================================================================

#87  Cox Regression
    file: 087_cox_PROC_PHREG_seedling_death_env.xlsx
  >> SCENARIO (narration):
    Here we model the risk of seedling death as a function of environmental
    conditions and seed source. A Cox proportional hazards regression relates
    the hazard over duration_month, censored by death_event, to elevation_m,
    precipitation_mm, dbh_baseline_cm and provenance origin. Cox regression
    suits this question because it estimates how each covariate scales mortality
    risk without specifying the baseline hazard shape. We learn which
    provenances and which site conditions carry elevated death risk for
    outplanted seedlings.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death_event
    - Predictor: provenance
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Predictor: dbh_baseline_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 102 (72.9%)   Concordance = 0.561
    elevation: HR = 1.0005 p = 0.020 *   (with precipitation, baseline-diameter covariates)

>> COMMENTARY (narration):
    We modelled seedlings' death risk from environmental variables (elevation, precipitation, baseline diameter) with
    Cox regression. Elevation is significant (HR = 1.0005, p = 0.020): death risk rises slightly with altitude
    (high-altitude stress). Concordance 0.56 means the model discriminates modestly -- seedling death is partly
    explained by environmental variables, with other factors also at play. Cox regression is the gold-standard
    survival model relating "time-to-event" to several predictors; it gives hazard ratios (HR). Relating seedling
    death risk to environmental conditions is valuable for afforestation site selection and choosing climate-resilient
    provenances.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT)
    file: 088_parametrik_AFT_PROC_LIFEREG_Weibull.xlsx
  >> SCENARIO (narration):
    In this study we model the actual survival time of seedlings to understand
    how drought stress and origin lengthen or shorten establishment time. Using
    a parametric accelerated failure time model with a Weibull distribution, we
    relate duration_month and the death_event indicator to drought_score,
    elevation_m and provenance. The AFT framework directly expresses how
    covariates accelerate or decelerate time-to-death, which is intuitive for
    forecasting survival duration. This tells us how much drought exposure
    shortens expected seedling lifespan across provenances.
  >> VARIABLE SELECTION:
    - Time: duration_month
    - Event: death_event
    - Predictor: provenance
    - Predictor: elevation_m
    - Predictor: drought_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution = Weibull   Median survival = 32.25 months
    Survival ~ elevation + drought score

>> COMMENTARY (narration):
    We analysed seedling survival time with a parametric Weibull model. While Cox is semi-parametric, Weibull models
    the whole survival curve with a specific mathematical distribution -- if the data fit it, it gives more efficient
    estimates and allows extrapolation. Median survival is 32.25 months. The AFT (accelerated failure time)
    interpretation is intuitive: how many times a factor like elevation/drought lengthens/shortens the survival time.
    When the time distribution is known, parametric models give more powerful and predictable results than Cox.
    Modelling seedling survival time is valuable for afforestation planning and cost forecasting.

====================================================================================

#89  Competing Risks
    file: 089_competing_risks_seedling_3sonuc.xlsx
  >> SCENARIO (narration):
    We study seedling fate when several distinct outcomes compete with one
    another over time. A competing risks analysis follows seedlings over
    duration_year, where the result variable distinguishes censoring, death,
    transplant and disease, and we examine how drought_score and provenance
    influence each outcome. Standard survival methods would mishandle these
    mutually exclusive endpoints, so a competing risks framework correctly
    partitions the cumulative incidence of each fate. This reveals, for example,
    whether drought drives death specifically rather than transplant or disease.
  >> VARIABLE SELECTION:
    - Time: duration_year
    - Event: result_0_censored_1_death_2_transplant_3_disease
    - Predictor: provenance
    - Predictor: drought_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 38   |   3 different outcomes (1=death, 2=transplant, 3=disease)
    Separate cumulative incidence (CIF) and risk for each event type

>> COMMENTARY (narration):
    A seedling can meet more than one outcome: death, transplant (uprooted and replanted elsewhere), or disease -- and
    these are mutually exclusive (once one happens the others cannot). Classic survival handles this incorrectly;
    competing-risks analysis produces a separate cumulative incidence for each outcome type. In the data 38 seedlings
    are still healthy (censored), the rest met different fates. This framework correctly answers questions like "which
    outcome's risk does drought increase". In forest-nursery trials, when there are different outcome pathways
    (death/transplant/disease), competing-risks analysis is the correct and informative method -- otherwise one
    outcome's risk is overstated.

====================================================================================

#90  Time-Dependent Cox
    file: 090_time_dep_cox_dbh_time_variable.xlsx
  >> SCENARIO (narration):
    This analysis asks whether a seedling's changing size over time affects its
    instantaneous mortality risk. Using a time-dependent Cox model with start
    and stop interval coding, we relate the event indicator to the time-varying
    dbh_t together with drought_score. Because diameter grows throughout the
    study, treating dbh_t as a covariate that updates within each interval
    captures its real-time effect on hazard. The model shows whether seedlings
    that are currently thinner face higher death risk at any given moment.
  >> VARIABLE SELECTION:
    - Start: start
    - Stop: stop
    - Event: event
    - Time-varying predictor: dbh_t
    - Predictor: drought_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    dbh_t: HR = 0.516   95% CI (0.389 , 0.685)   p < .001 ***
    Time-varying stem diameter + drought covariates

>> COMMENTARY (narration):
    A seedling's stem diameter grows over time; standard Cox assumes it constant, time-dependent Cox uses the current
    diameter in each time interval (start-stop format). Current diameter is highly significant and protective (HR =
    0.516, p < .001): seedlings whose diameter is large at a given moment have half the death risk -- large/healthy
    seedlings are more stress-resistant. This is a critical inference: the current growth status is a better survival
    indicator than the baseline diameter. Time-dependent Cox is the correct method when "the current, continuously
    monitored condition matters, not just the baseline"; it is the standard for modelling time-varying growth
    measurements in longitudinal seedling monitoring.

====================================================================================

#91  Survey-Weighted Cox
    file: 091_survey_phreg_complex_survey_survival.xlsx
  >> SCENARIO (narration):
    Our forest genetics team monitored seedling survival across a complex multi-
    stage sampling design, where seedlings were grouped into strata and sampling
    clusters with unequal selection probabilities. We want to know how survival
    time (duration_year) until death (death) is shaped by provenance and
    elevation_m, while correctly accounting for the stratum, cluster_id, and
    weight structure of the survey. A survey-weighted Cox proportional hazards
    model is the right tool because it estimates hazard ratios for provenance
    and elevation while respecting the stratified, clustered, weighted design so
    that population-level inference is unbiased. This lets us compare drought-
    adapted seed sources without the design itself distorting our survival
    estimates.
  >> VARIABLE SELECTION:
    - Stratum: stratum
    - Cluster: cluster_id
    - Weight: weight
    - Predictor: provenance
    - Time: duration_year
    - Event: death
    - Covariate: elevation_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.49   Weighted/clustered survival (elevation covariate)
    Death ~ elevation

>> COMMENTARY (narration):
    We combined survival analysis with a complex sample design: seedlings come from a population sampled with weights
    and clusters. Survey-PHREG runs the Cox model with design weights and cluster-robust standard errors, so the
    hazard ratios and confidence intervals generalise to the population. Concordance 0.49 means elevation alone is a
    weak discriminator in this model. When national/regional forest-monitoring data have "time-to-event" data,
    ignoring the design biases the estimates; survey-PHREG provides valid inference.

====================================================================================

#92  Interval-Censored Survival
    file: 092_interval_censored_seedling_death_interval.xlsx
  >> SCENARIO (narration):
    In a nursery trial we recorded seedling mortality only at periodic
    inspections, so for each seedling we know that death occurred somewhere
    between a lower_bound_month and an upper_bound_month rather than at an exact
    day. We ask whether time-to-death differs among provenance seed sources when
    the exact event time is unknown. Interval-censored survival analysis is
    appropriate here because it models the survival function directly from the
    lower and upper bounding intervals instead of pretending we observed precise
    death times. This gives us honest provenance survival comparisons despite
    the coarse inspection schedule.
  >> VARIABLE SELECTION:
    - Lower bound: lower_bound_month
    - Upper bound: upper_bound_month
    - Group: provenance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 69   Median survival = 116.0 months   (Turnbull algorithm)
    Death time known within [lower, upper] interval

>> COMMENTARY (narration):
    In forest monitoring we rarely know the exact time of an event (seedling death): we visit the trial periodically,
    find a seedling alive at one visit and dead at the next -- death happened somewhere between the two dates.
    Interval-censored survival (Turnbull algorithm) handles exactly this uncertainty, avoiding fixing the event to an
    arbitrary date. Median survival is 116 months. By the nature of periodic inspection (annual/periodic measurement),
    event times are always within an interval; forcing this data to a single point biases the result. The interval-
    censored method matches the reality of periodic forest-monitoring data.

====================================================================================

#93  Frailty Cox
    file: 093_frailty_cox_provenance_random_frailty.xlsx
  >> SCENARIO (narration):
    We followed seedlings drawn from several provenances and suspect that trees
    sharing the same seed source experience correlated, source-specific survival
    risks beyond their measured traits. We model time-to-death (duration_year,
    death) using elevation_m and drought_score as fixed covariates, while
    treating provenance as a random frailty term. A frailty Cox model fits
    because it adds a provenance-level random effect (frailty) to the
    proportional hazards structure, capturing unobserved shared susceptibility
    among seedlings from the same provenance. This separates genuine drought and
    elevation effects from latent between-provenance heterogeneity in survival.
  >> VARIABLE SELECTION:
    - Frailty: provenance_frailty
    - Covariate: elevation_m
    - Covariate: drought_score
    - Time: duration_year
    - Event: death

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.557   Provenance-clustered shared frailty (random effect)
    Death ~ elevation + drought  +  (provenance frailty)

>> COMMENTARY (narration):
    Seedlings are clustered within provenances; seedlings of the same provenance share unmeasured common genetic
    factors and carry similar death risk. Frailty Cox adds a shared "frailty" (random effect) for each provenance to
    model this clustering -- the mixed-model version of survival analysis. Concordance 0.56. Accounting for
    provenance-level hidden genetic differences gives both correct standard errors and information on "how much genetic
    heterogeneity exists in survival among provenances" -- which points to the heritability of survival. It is the
    correct method for clustered forest-survival data (provenances, families).

====================================================================================

#94  Time Series Description
    file: 094_time_series_25yil_clone_growth_trial.xlsx
  >> SCENARIO (narration):
    A clonal growth trial was measured once per year over 25 consecutive years,
    recording mean_dbh_cm and mean_height_m alongside the climate of each year.
    We want to describe how stand growth and its climatic context
    (annual_precipitation_mm, mean_temperature_C) evolved across the time series
    indexed by year. A time series description is the natural first step because
    it summarizes level, trend, and year-to-year variation in the growth and
    climate series before any modeling. This gives us a clear narrative of the
    trial's long-term growth trajectory.
  >> VARIABLE SELECTION:
    - Time: year
    - Series: mean_dbh_cm
    - Series: mean_height_m
    - Series: annual_precipitation_mm
    - Series: mean_temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    25 observations (yearly)   ADF p = 0.988 -> NOT stationary   Trend = increasing
    25-year clone growth trial (mean dbh)

>> COMMENTARY (narration):
    We examined the mean stem-diameter (dbh) series of a 25-year clone growth trial. The ADF test does not find it
    stationary (p = 0.988) -- as expected there is a strong increasing trend (trees grow over the years). This
    pre-diagnosis is critical: most time-series models (like ARIMA) require stationarity, so differencing/
    transformation may be needed first. Time-series analysis reveals the structure of a long-term growth trial --
    trend, stationarity. In forest genetics, long-running (multi-decade) growth trials are the basis of tracking
    clone/provenance performance over time; this analysis quantifies that structure.

====================================================================================

#95  STL Decomposition
    file: 095_stl_ayrisma_chlorophyll_monthly_mevsimsel.xlsx
  >> SCENARIO (narration):
    We logged monthly chlorophyll measurements (chlorophyll_unit) over twelve
    years, indexed by date with its year and month, and expect a strong seasonal
    photosynthetic rhythm. The question is how much of the chlorophyll signal is
    seasonal versus underlying trend versus irregular noise. STL decomposition
    is the right tool because it separates the monthly chlorophyll_unit series
    into seasonal, trend, and remainder components using locally weighted
    regression. This reveals the recurring phenological cycle and any slow drift
    in foliar chlorophyll over the years.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: chlorophyll_unit

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Period (seasonal length) = 12 (monthly -> annual cycle)
    Series decomposed into trend + seasonal + residual components

>> COMMENTARY (narration):
    STL (Seasonal-Trend decomposition using Loess) splits a monthly chlorophyll series into three components:
    long-term trend, a recurring seasonal cycle (12-month = annual) and the remaining residual. The annual period is
    meaningful -- chlorophyll (photosynthetic activity) fluctuates seasonally (high in summer, low in winter; leaf
    drop/seasonal activity). This decomposition clarifies "is chlorophyll generally rising, or just fluctuating each
    year". STL is the most intuitive way to interpret seasonal physiological series (chlorophyll, photosynthesis,
    phenology), separating trend from season to see the tree's long-term physiological tendency.

====================================================================================

#96  ARIMA
    file: 096_arima_bud_flushing_monthly.xlsx
  >> SCENARIO (narration):
    We tracked a monthly bud_flushing_score over fifteen years, indexed by date,
    to study the phenology of bud break in our provenance trial. We want to
    model the temporal dependence in this series and forecast future bud-
    flushing behavior. An ARIMA model fits because it captures autocorrelation,
    differencing for trend, and moving-average structure in the
    bud_flushing_score series to produce statistically grounded forecasts. This
    helps anticipate shifts in flushing timing under changing conditions.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: bud_flushing_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model fitted   AIC = 1209.22
    Monthly bud-flushing score -- future forecast

>> COMMENTARY (narration):
    We fitted an ARIMA model to a monthly bud-flushing score series and forecast the future. ARIMA combines the
    series' own past values (AR), past errors (MA) and differencing (I, to make it stationary). With AIC = 1209 the
    best model was selected. Forecasts rest on the observed trend and seasonal autocorrelation structure and come with
    an uncertainty band. Phenology forecasting is critical in climate-change studies: anticipating future bud-flush
    timing enables frost-risk and phenological-mismatch assessment. ARIMA is the classic forecasting method for
    phenological series with seasonality and trend.

====================================================================================

#97  Exponential Smoothing
    file: 097_ets_holt_winters_growth_monthly.xlsx
  >> SCENARIO (narration):
    Monthly growth measurements (growth_unit) were collected over ten years,
    indexed by date, and show both a developing trend and a repeating seasonal
    pattern. We want a forecasting method that adapts to evolving level, trend,
    and seasonality in seedling growth. Exponential smoothing (Holt-Winters) is
    appropriate because it weights recent observations more heavily and
    explicitly models trend and seasonal components of the growth_unit series.
    This yields smooth, responsive forecasts of upcoming growth.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: growth_unit

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method = Holt-Winters (additive seasonal, period 12)   AIC = 276.00
    Monthly growth-unit forecast

>> COMMENTARY (narration):
    Holt-Winters exponential smoothing estimates the level + trend + seasonality components in a weighted way (giving
    more weight to the recent past). It captured the seasonal pattern of the monthly growth series (AIC = 276, low --
    good fit). Tree growth typically shows strong seasonality (fast in spring-summer, dormant in autumn-winter);
    Holt-Winters is often very successful for such series and is more intuitive than ARIMA. In forest growth
    forecasting it is an easy-to-build, interpretable alternative; it is valuable for anticipating seasonal growth
    dynamics.

====================================================================================

#98  Mann-Kendall Trend
    file: 098_mann_kendall_bud_fracture_40yil_trend.xlsx
  >> SCENARIO (narration):
    We assembled a 40-year record of bud fracture timing (bud_fracture_jd, in
    Julian days) together with annual_precipitation_mm and mean_temperature_C,
    indexed by year. We ask whether bud fracture phenology has shifted
    monotonically over four decades, possibly tracking climate change. The Mann-
    Kendall trend test fits because it is a nonparametric test for a monotonic
    trend that does not assume normality, ideal for detecting directional change
    in bud_fracture_jd and the climate series. This lets us state whether
    phenology and climate are trending significantly across the years.
  >> VARIABLE SELECTION:
    - Time: year
    - Series: bud_fracture_jd
    - Series: annual_precipitation_mm
    - Series: mean_temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p = 0.0014 **   Trend = DECREASING
    Bud-flush day 40-year trend

>> COMMENTARY (narration):
    We tested whether the bud-flush day (Julian day, phenology) shows a significant long-term trend over 40 years with
    the Mann-Kendall test -- a nonparametric, outlier- and distribution-robust trend test. The result is significant
    (p = 0.001) and the direction is decreasing: the bud-flush day falls consistently over the years (shifts earlier).
    This is strong evidence that climate warming pulls tree phenology earlier each decade -- spring arrives sooner,
    trees wake earlier. Mann-Kendall + Sen's slope is the gold standard for detecting long-term phenological/climate
    trends; it is one of the most sensitive indicators documenting tree phenology's response to climate change.

====================================================================================

#99  Anomaly Detection
    file: 099_anomali_tespiti_monthly_growth_outlier.xlsx
  >> SCENARIO (narration):
    We recorded a monthly mean_height_growth_cm series over fifteen years,
    indexed by date, and want to flag months where growth departed sharply from
    the expected pattern, possibly due to drought, pests, or measurement error.
    The goal is to identify unusual observations rather than describe the whole
    series. Anomaly detection is the right approach because it models the normal
    behavior of the mean_height_growth_cm series and isolates statistically
    extreme points as outliers. This surfaces the specific months warranting
    closer ecological investigation.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: mean_height_growth_cm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 7   (IQR-based outlier detection)
    Extraordinary values in the monthly height-growth series

>> COMMENTARY (narration):
    We detected the extraordinary values (anomalies) in the monthly height-growth series -- 7 extraordinary months
    were found. These anomalies can point to real events: an extraordinarily good growth month (ideal precipitation/
    temperature), a stress event (growth halt due to drought, frost, insect attack), or measurement error. Anomaly
    detection automatically answers "when did something go differently" in continuous growth-monitoring data and is
    the basis of early-warning systems. In forest monitoring it is extremely practical for picking out the few
    critical growth events that require attention from among many measurements.

====================================================================================

#100  Variance Components (VARCOMP)
    file: 100_varcomp_provenance_family_h2_NESTED.xlsx
  >> SCENARIO (narration):
    In a nested provenance-family trial we measured 5-year height
    (height_cm_5yr) on individual trees, where families are nested within
    provenances and trees within families. We want to partition the total
    phenotypic variance into provenance, family, and tree levels in order to
    estimate genetic structure and heritability. A variance components (VARCOMP)
    analysis with this nested random structure is appropriate because it
    apportions variation in height_cm_5yr among provenance_random,
    family_id_random, and tree_random effects. The resulting variance ratios let
    us quantify how much of the growth variation is genetically structured.
  >> VARIABLE SELECTION:
    - Random (level 1): provenance_random
    - Random (nested): family_id_random
    - Random (nested): tree_random
    - Response: height_cm_5yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    provenance = 63.2%   family (nested in provenance) = 14.2%   residual = 22.6%
    Nested: provenance > family > tree (REML)

>> COMMENTARY (narration):
    We partitioned the variability of 5-year height by which genetic level -- provenance, family or individual -- it
    comes from with variance-components analysis: provenance and family are nested. The result is striking: 63% of the
    variance comes from between-provenance differences, 14% from (within-provenance) between-family differences, 23%
    from individual/residual variation. This is the result at the HEART of forest genetics: most of the genetic
    structure is at the provenance level -- i.e. "where the seed comes from" is the biggest genetic factor determining
    height. The family level also contributes significantly (heritability is present). These variance components are
    the basis for calculating narrow-/broad-sense heritability (h2) and genetic gain -- the scientific foundation of
    breeding strategy (provenance vs family selection) and seed-transfer zones.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_best_control_irrigated_BF10.xlsx
  >> SCENARIO (narration):
    In a provenance trial we wanted to know whether supplemental irrigation
    improves early seedling growth in a drought-prone forest nursery. We
    measured height_cm_3yr on 60 seedlings split between a control group and an
    irrigated treatment. Because we want to quantify the evidence for and
    against an irrigation effect rather than just reject a null, we use a
    Bayesian t-Test, which gives us a Bayes factor (BF10) comparing the two
    treatment means directly. This tells us how strongly the data support a real
    height difference between control and irrigated seedlings.
  >> VARIABLE SELECTION:
    - Group: treatment
    - Dependent: height_cm_3yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 2.28e+40 (decisive evidence)   Cohen's d = 5.05
    Irrigated vs control -- 3-year height

>> COMMENTARY (narration):
    We compared the 3-year height of irrigated vs control (non-irrigated) seedlings with a Bayesian t-test. The Bayes
    Factor BF10 = 2.28e+40 -- astronomically large, "decisive evidence": the data support the hypothesis of a
    difference quadrillions of times more than no difference. Cohen's d = 5.05 makes the effect extraordinarily large.
    Irrigation increased seedling growth overwhelmingly. Unlike a classic p-value, the Bayes Factor directly measures
    the strength of evidence for both H1 and H0 and distinguishes "absence of evidence" from "evidence of absence".
    Irrigation's effect is not just statistical but practically overwhelming -- the Bayesian framework shows this
    convincingly.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_midparent_offspring_h2.xlsx
  >> SCENARIO (narration):
    We are estimating narrow-sense heritability of stem diameter in a forest
    tree breeding program by relating parents to their offspring. For 50
    families we recorded midparent_dbh_cm and the corresponding
    mean_offspring_dbh. A Bayesian Correlation lets us estimate the parent-
    offspring association along with full credible intervals, which is ideal
    when we want to express uncertainty in a heritability-related slope. This
    Bayesian approach naturally handles our modest sample of families while
    quantifying how reliably parental diameter predicts offspring diameter.
  >> VARIABLE SELECTION:
    - Variable 1: midparent_dbh_cm
    - Variable 2: mean_offspring_dbh

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.750 (very strong)   BF10 = 3.32e+07 (decisive evidence)
    Midparent dbh - offspring dbh relationship

>> COMMENTARY (narration):
    We assessed the relationship between midparent diameter and offspring diameter with Bayesian correlation: r = 0.75
    is very strong and the Bayes Factor BF10 = 3.32e+07 makes the evidence for the relationship's existence decisive.
    This is a classic heritability (h2) setup: the parent-offspring correlation shows how much of the trait is
    heritable -- a high correlation means high heritability (diameter is largely under genetic control). Bayesian
    correlation expresses the strength of the relationship with a probability distribution and an evidence factor.
    Heritability is the foundation of tree breeding: high h2 shows rapid genetic gain is achievable by selection. This
    result proves diameter is highly amenable to breeding.

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_provenance_site_2yonlu.xlsx
  >> SCENARIO (narration):
    We tested whether seed provenance and planting site jointly shape juvenile
    growth across a multi-site common-garden experiment. The dataset holds
    height_cm_3yr for 180 trees from four provenances (Antalya, Mugla, Mersin,
    Burdur) grown at three sites (Ankara, Eskisehir, Konya). A two-way Bayesian
    ANOVA estimates the provenance and site main effects plus their interaction,
    giving posterior distributions instead of single p-values. This helps us
    judge whether some provenances are broadly superior or whether their ranking
    changes from site to site (genotype-by-environment interaction).
  >> VARIABLE SELECTION:
    - Factor 1: provenance
    - Factor 2: site
    - Dependent: height_cm_3yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    provenance: BF10 = 6.23e+19 (decisive evidence)
    Provenance x site -> 3-year height

>> COMMENTARY (narration):
    We examined the effect of provenance and site factors on height with Bayesian ANOVA. The provenance effect is
    overwhelming: Bayes Factor 6.23e+19 -- "decisive evidence". So different provenances differ markedly in height
    (strong genetic differentiation among provenances). Bayesian ANOVA gives a separate Bayes Factor for each effect
    and interaction, answering "which factor really matters" more informatively than classic ANOVA -- and can even
    provide "evidence of absence" for non-significant effects. In provenance trials (provenance x site) it is a
    powerful way to evaluate factor effects with evidence; it provides the Bayesian evidential basis for seed-transfer
    decisions.

====================================================================================

#104  Bayesian Hierarchical
    file: 104_hierarchical_bayesian_3seviye_region_provenance_family.xlsx
  >> SCENARIO (narration):
    Our growth data are naturally nested: families sit within provenances, which
    sit within broader geographic regions. For 240 trees we recorded
    height_cm_5yr alongside region (mediterranean, Toros, black_sea),
    provenance, and family_id. A three-level Bayesian Hierarchical model
    partitions variance across region, provenance, and family while sharing
    information across these levels (partial pooling). This is the right tool
    because it gives us robust variance components at each genetic level even
    when some families have few trees.
  >> VARIABLE SELECTION:
    - Level 3: region
    - Level 2: provenance
    - Level 1: family_id
    - Dependent: height_cm_5yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group (region) random variance = 72.84 (61.2%)   residual = 46.10 (38.8%)   ICC = 0.612
    Height ~ provenance  +  (1 | region)   (Bayesian)

>> COMMENTARY (narration):
    We analysed nested genetic data (region > provenance > family > tree) with a Bayesian hierarchical model -- the
    Bayesian counterpart of LMM. The region-level variance is 61% of the total (ICC = 0.612); most of the variability
    comes from between-region differences. So height depends strongly on which region one comes from (geographic-
    genetic structure). The Bayesian hierarchical model's strength: it expresses both within- and between-group
    uncertainty with full probability distributions and balances groups with few observations via "partial pooling" --
    very valuable in few-family forest trials. It is the modern way to model multi-level genetic structure (region/
    provenance/family).

====================================================================================

#105  Spatial Lag (SAR)
    file: 105_spatial_sar_population_IBD_He.xlsx
  >> SCENARIO (narration):
    We investigated isolation-by-distance in genetic diversity, asking whether
    nearby populations have similar expected heterozygosity because of gene
    flow. For 100 populations we have He_mean together with their lat and lon
    coordinates and environmental drivers elevation_m, precipitation_mm, and
    temperature_C. A Spatial Lag (SAR) model adds a spatially lagged response
    term so we can test whether a population's diversity depends on its
    neighbors' diversity beyond what climate explains. This separates true
    spatial spillover in He_mean from environmental effects.
  >> VARIABLE SELECTION:
    - Dependent: He_mean
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Predictor: temperature_C
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho (spatial lag) = 0.76   z = 12.11   p < .001   Pseudo R2 = 0.62
    He ~ elevation + precipitation + temperature  +  neighbour He

>> COMMENTARY (narration):
    Modelling populations' heterozygosity (He), we accounted for spatial dependence -- the tendency of nearby
    populations to be similar. The SAR (spatial autoregressive) model links a population's He to its neighbours' He
    too (rho). Here rho = 0.76, very strongly positive (z = 12.11, p < .001): neighbouring populations' genetic
    diversity is very similar. This is the fundamental population-genetics pattern of IBD -- "Isolation By Distance":
    nearby populations are similar through gene flow, distant ones differentiate. Using a spatial model is critical,
    because ordinary regression gives spurious results when spatial autocorrelation is present. SAR is the right way
    to model the gene-flow and IBD process.

====================================================================================

#106  Spatial Error
    file: 106_spatial_error_gene_akisi_residual_autocorr.xlsx
  >> SCENARIO (narration):
    We modeled how climate and stand age structure relate to genetic diversity
    while accounting for spatially correlated residuals left by unmeasured gene
    flow. Across 90 populations we recorded He_mean, anavar_precipitation_mm,
    and age_structuring_year, with lat and lon for spatial structure. A Spatial
    Error model places the spatial dependence in the error term, which is
    appropriate when omitted spatial processes, not neighbor diversity itself,
    drive the autocorrelation. This yields unbiased estimates of the
    precipitation and age effects on He_mean.
  >> VARIABLE SELECTION:
    - Dependent: He_mean
    - Predictor: anavar_precipitation_mm
    - Predictor: age_structuring_year
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda (spatial error) = 0.50   z = 4.45   p < .001   Pseudo R2 = 0.08
    He ~ environmental variables  +  spatial error structure

>> COMMENTARY (narration):
    The spatial error model addresses a different kind of spatial dependence than SAR: not a direct effect of
    neighbour values, but correlation induced through the residuals by unobserved (omitted) spatial factors. Lambda =
    0.50 (z = 4.45, p < .001) shows this spatial error structure is strong -- there are unmeasured common spatial
    factors (gene-flow corridors, historical range, unobserved environment). Which spatial model is appropriate (SAR
    or SEM) depends on whether the effect comes from neighbour values or from unmeasured factors. SEM cleans out these
    unmeasured spatial confounders to estimate the true effect of environmental variables on He more accurately.

====================================================================================

#107  GWR
    file: 107_gwr_local_He_elevation_relationship.xlsx
  >> SCENARIO (narration):
    We suspected that the relationship between elevation and genetic diversity
    is not constant but varies geographically across a mountainous range. For
    110 populations we have He_mean, elevation_m, and precipitation_mm, each
    tied to lat and lon. Geographically Weighted Regression (GWR) fits local
    regressions around each location, producing a map of how strongly
    elevation_m predicts He_mean in different areas. This reveals spatial non-
    stationarity that a single global model would mask.
  >> VARIABLE SELECTION:
    - Dependent: He_mean
    - Predictor: elevation_m
    - Predictor: precipitation_mm
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.276   (local/location-specific regression)
    Spatial variation of the He-elevation relationship

>> COMMENTARY (narration):
    Ordinary regression assumes a single He-elevation relationship for the whole region; but this relationship can
    vary in space -- in one region elevation may strongly affect genetic diversity while in another it is weak. GWR
    (Geographically Weighted Regression) captures this spatial heterogeneity by fitting a separate local regression at
    each location (R2 = 0.28). The output is a map of coefficients: a surface showing where elevation's effect on He
    is strong and where it is weak. This tests "is the relationship the same everywhere" and sets local conservation
    priorities. In landscape genetics, GWR reveals what the global model hides for mapping the spatial variation of
    the environment-gene relationship.

====================================================================================

#108  Nested LMM
    file: 108_nested_lmm_R_P_F_P_agnostik.xlsx
  >> SCENARIO (narration):
    In a field trial laid out in blocks, families are nested within populations
    and we want clean variance components for early growth. The data give
    height_cm_5yr for 360 trees with block (R1-R4), population (P1_Antalya
    through P5_KMaras), and family_no (F01-F06) nested inside population. A
    Nested LMM treats block as a design factor and models family within
    population as nested random effects. This correctly attributes growth
    variation to populations versus families and avoids confounding the two
    genetic levels.
  >> VARIABLE SELECTION:
    - Fixed: block
    - Random (nesting): population
    - Random (nested): family_no
    - Dependent: height_cm_5yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    population (P): F = 58.84, p < .001 ***   |   family x population (RxP): F = 13.05, p < .001 ***
    family [within population] (R): F = 10.38, p < .001 ***

>> COMMENTARY (narration):
    In this design the measurements are nested: population > family > tree. The nested mixed model tests each level's
    contribution to the variance separately. The results are significant at all levels: between-population differences
    are largest (F = 58.84), but there are real differences at the family and family×population levels too. This is
    the classic nested structure of forest-genetics trials (population/family/individual) and is the basis for
    separating heritability components. Mixing up the levels in nested data (e.g. ignoring population) produces
    spurious significance or wrong standard errors. Nested LMM is the correct framework for hierarchical genetic
    designs, separating the genuine contribution of each scale.

====================================================================================

#109  Crossed LMM
    file: 109_crossed_lmm_clone_site_diallel_tipli.xlsx
  >> SCENARIO (narration):
    We evaluated clone performance across multiple test sites where every clone
    is replicated at each location, a fully crossed layout. The dataset records
    height_cm_5yr for 200 entries with clone and site (S1-S4) crossed rather
    than nested. A Crossed LMM fits clone and site as crossed random effects,
    letting us estimate clone genetic variance and site variance simultaneously.
    This is the right model for detecting whether clone rankings are stable
    across sites or shift with the environment.
  >> VARIABLE SELECTION:
    - Random (crossed): clone
    - Random (crossed): site
    - Dependent: height_cm_5yr

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    clone (A): F = 8.75, p < .001 ***   |   site (B): F = 0.84, p = 0.482 ns
    clone x site interaction: F = 11.35, p < .001 *** (SIGNIFICANT -- strong G×E)

>> COMMENTARY (narration):
    Unlike nesting, here two factors are crossed: each clone was tested at each site (fully factorial). A striking and
    very important result: the clone main effect is significant (F = 8.75), the site main effect non-significant
    (F = 0.84), BUT the clone×site interaction is highly significant (F = 11.35, p < .001). This is a strong
    Genotype×Environment (G×E) interaction: the ranking of clones changes from site to site -- the best clone at one
    site may not be best at another. This is one of the most critical inferences in forest genetics: "there is no
    single superior clone for all sites"; clone selection must be site-specific. The crossed mixed model is the right
    way to capture G×E interaction and is the basis of clonal forestry / seed-transfer policy.

====================================================================================

#110  KDE Density Map
    file: 110_KDE_population_sample_density.xlsx
  >> SCENARIO (narration):
    We wanted to visualize where genetically sampled populations cluster most
    densely across the landscape. For 95 sampling units we have lat and lon
    coordinates along with He_mean. A Kernel Density Estimation (KDE) density
    map smooths these point locations into a continuous surface of sampling
    intensity. This highlights spatial gaps and hotspots in our diversity
    sampling effort so we can plan future collections.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Weight: He_mean

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    95 population points   mean He = 0.404
    Sampling-density map via kernel density estimation

>> COMMENTARY (narration):
    On the Map tab, we render the geographic density of population sampling points into a heat map with KDE (Kernel
    Density Estimation). KDE turns points into a smooth density surface: we can see where sampling/populations are
    dense and where there are gaps. The distribution of 95 populations shows how genetic sampling effort concentrates
    geographically. This is used both to detect sampling gaps (which regions are under-represented -- future sampling
    priority) and to see high-population-density regions. It is the basic mapping tool for visualising spatial genetic
    data; it is valuable in gene-conservation area planning.

====================================================================================

#111  Hexbin Map
    file: 111_Hexbin_provenance_distribution.xlsx
  >> SCENARIO (narration):
    In this study we map the geographic distribution of forest seed provenances
    across our survey region to understand where reproductive material is
    concentrated. For each sampling unit we recorded its position as lat and
    lon, together with a density_interval class that flags whether stocking is
    low, medium or high. By aggregating these points into a hexagonal binning
    grid we can see the spatial intensity of provenances without the clutter of
    thousands of overlapping markers. A Hexbin Map is the right tool here
    because it summarises point density over equal-area hexagons, revealing core
    provenance zones and sparsely sampled gaps at a glance.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Category: density_interval

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    150 provenance/unit points   aggregated into hexagonal cells
    Each cell shows the point density within it by colour

>> COMMENTARY (narration):
    The hexbin map aggregates many points into hexagon cells and shows each cell's density by colour -- a clear
    solution, alternative to KDE, when thousands of points overlap into an unreadable mess. We rendered 150 provenance
    points into a hexagonal grid to visualise where provenance density is highest. Hexagonal cells carry less
    edge-bias than squares and represent neighbour relations more evenly. It is the practical way to turn dense
    spatial genetic data (provenance locations, sampling points) into a readable density map.

====================================================================================

#112  Moran's I
    file: 112_Morans_I_He_IBD_autocorrelation.xlsx
  >> SCENARIO (narration):
    Here we ask whether expected heterozygosity is spatially structured across
    our provenance network, testing the classic isolation-by-distance
    expectation that nearby stands are genetically more similar than distant
    ones. Each unit carries its lat and lon coordinates and a He_mean value
    summarising its genetic diversity. Moran's I lets us quantify global spatial
    autocorrelation in He_mean, telling us whether high-diversity stands cluster
    together or are randomly scattered. This test fits because a single,
    interpretable statistic captures whether genetic diversity follows a
    geographic gradient, which directly informs how seed zones should be
    delineated.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Value: He_mean

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.239   z = 4.37   p < .001
    Spatial autocorrelation of heterozygosity (He) (positive -- IBD)

>> COMMENTARY (narration):
    Moran's I is the global spatial-autocorrelation statistic measuring "do nearby populations have similar genetic
    diversity". The result is significantly positive (I = 0.24, z = 4.37, p < .001): He is spatially clustered --
    high-diversity populations lie near each other, and so do low ones. This is the signature of the fundamental
    population-genetics pattern Isolation By Distance (IBD): gene flow makes nearby populations similar, and
    differentiation increases with distance. This finding is also methodologically important: if spatial
    autocorrelation exists, ordinary statistics are wrong (hence the spatial models in #105-107). Moran's I is the
    starting diagnostic of landscape genetics.

====================================================================================

#113  Getis-Ord Hotspot
    file: 113_Getis_Ord_high_He_hotspot.xlsx
  >> SCENARIO (narration):
    We want to pinpoint exactly where the most genetically diverse populations
    occur so that high-value seed sources can be prioritised for conservation.
    Using each unit's lat and lon position and its He_mean diversity score, we
    look for statistically significant local concentrations of high
    heterozygosity. The Getis-Ord Gi* hotspot analysis is ideal because it
    identifies clusters of high He_mean (hotspots) and low He_mean (coldspots)
    with a local significance test, rather than just describing overall pattern.
    The resulting hotspot map gives managers concrete locations of genetic
    reservoirs worth protecting.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon
    - Value: He_mean

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    29 hot spots (high-He cluster)   26 cold spots (low-He cluster)   n = 68
    Getis-Ord Gi* local statistic (z > 1.96 significant)

>> COMMENTARY (narration):
    While Moran's I reports the overall clustering of the whole region, Getis-Ord Gi* shows WHERE the hot/cold spots
    are, location by location. 29 populations are statistically significant "hot spots" (high genetic diversity
    surrounded by high diversity -- genetic-diversity centres), and 26 populations are "cold spots" (low-diversity
    clusters -- bottlenecked/isolated areas). This gives direct guidance for gene conservation: hot spots should be
    protected as genetic-diversity reservoirs, while cold spots may need genetic rescue. Gi* hot-spot analysis is the
    standard spatial method for finding "where genetic diversity concentrates geographically" -- the basis of
    conservation prioritisation.

====================================================================================

#114  DBSCAN Spatial
    file: 114_DBSCAN_izole_populasyonlar.xlsx
  >> SCENARIO (narration):
    In this analysis we search for isolated forest populations that may be at
    risk of inbreeding and genetic drift due to their geographic separation. We
    use only the spatial coordinates, lat and lon, of each sampling unit to
    study how the populations group in space. DBSCAN clustering is well suited
    here because it discovers dense population clusters of arbitrary shape while
    explicitly flagging sparse, isolated points as noise. Those flagged outliers
    highlight the spatially isolated stands that warrant closer genetic
    monitoring and possible assisted gene flow.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   noise (isolated) = 8   n = 115
    Population locations clustered by density

>> COMMENTARY (narration):
    On the map we clustered populations' geographic locations by density with DBSCAN: four dense population groups
    were found, and 8 populations were flagged as "isolated/noise" points belonging to no cluster. DBSCAN's strength
    is that it does not require the number of clusters in advance and can capture irregularly shaped clusters -- and
    automatically separates sparse/distant populations. In conservation genetics this is extremely valuable: isolated
    (noise) populations usually have unique genetic heritage, are cut off from gene flow, and are at risk of
    bottleneck/inbreeding -- i.e. priority conservation candidates. The four clusters show populations connected by
    gene flow. This spatial clustering is a practical tool for defining conservation units and detecting isolated
    populations. With this we complete the full 114-analysis tour of the Forest Genetics package.

====================================================================================

#115  ?
    file: 115_seed_mescereleri_maps_analysis.xlsx
  >> SCENARIO (narration):
    This study maps Turkey's registered forest seed stands to visualise how
    certified seed sources are distributed across regions and elevations. For
    every stand we have its X and Y map coordinates, its elevation, its stand
    area, plus descriptive attributes such as region, conservation status and
    the species FORM. By plotting each seed stand as a point symbolised by
    region and sized or coloured by elevation and area, we can reveal whether
    seed sources are evenly spread or concentrated in particular landscapes. A
    thematic point map is the appropriate tool here because it places
    categorical and numeric stand attributes directly onto geographic space,
    supporting decisions about gaps in seed-source coverage.
  >> VARIABLE SELECTION:
    - Longitude: X
    - Latitude: Y
    - Value: elevation
    - Value: area
    - Category: region
    - Category: conservation
    - Category: FORM

  ----- RESULT & COMMENTARY -----

====================================================================================

#116  Mixed-Design (Split-Plot) ANOVA
    file: 116_mixed_anova_seedling_height_cm.xlsx
  >> SCENARIO (narration):
    We follow 40 seedlings measured at three year levels (year1, year2, year3); each belongs to one of two provenance groups (northern / southern). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on seedling height.
  >> VARIABLE SELECTION:
    - Dependent variable: seedling_height_cm
    - Subject ID: tree_id
    - Between-subjects factor: provenance
    - Within-subjects factor: year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (provenance): F(1,38) = 37.78  p < .001  np2 = 0.499
    Within (year):  F(2,76) = 659.51  p < .001  np2 = 0.946
    Interaction:         F(2,76) = 10.84  p < .001  np2 = 0.222
    Mauchly W = 0.671  p = 0.001   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a provenance x year mixed design we analyzed seedling height for 40 seedlings (120 observations). The interaction is significant (F(2,76) = 10.84, p < .001, np2 = 0.222) *** -- the two groups' change across year differs in magnitude. The between-subjects main effect (northern vs southern) is F = 37.78, p < .001; the within-subjects main effect (year1/year2/year3) is F = 659.51, p < .001. Mauchly's test p = 0.001, so sphericity is violated, so the Greenhouse-Geisser corrected within p is read. Read the interaction first: when it is significant the group effect must be interpreted separately at each year level. In forest genetics, the mixed design is the standard analysis for provenance trials tracking growth across years.

====================================================================================
