====================================================================================
MerQur - Spor_Bilimleri - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Spor_Bilimleri/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_athlete.xlsx
  >> SCENARIO (narration):
    We compiled the inventory of 280 athletes. Age, height, weight, VO2max, training
    hours and performance were measured for each. Before any inferential test we
    want the overall picture of the sample; so we begin with descriptive statistics.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: height_cm
    - Variables: weight_kg
    - Variables: VO2_max
    - Variables: training_hour_week
    - Variables: performance_score
    - Grouping (categorical): sport_branch

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 athletes
    age: mean = 25.3   |   height: mean = 177.7 cm   |   weight: mean = 74.8 kg
    VO2max: mean = 55.3   |   weekly training: mean = 10.1 hours   |   performance: mean = 75.7

>> COMMENTARY (narration):
    First we draw the overall picture of our athlete sample: 280 athletes, average age ~25, height 178 cm, weight
    75 kg, VO2max 55.3, 10 weekly training hours and a performance score of 75.7. This descriptive table lays the
    groundwork for every analysis that follows -- discipline/level comparisons, physiology-performance relationships,
    spatial performance pattern. In sports science, before any inferential test, summarizing the athlete sample's basic
    features (age, anthropometry, VO2max, training, performance) is essential both to audit data quality and to set
    study priorities.

====================================================================================

#2  Normality Tests
    file: 02_normality_sprint_VO2_lactate.xlsx
  >> SCENARIO (narration):
    We examine whether 100m sprint time, VO2max and lactate are normally
    distributed. Because subsequent t-tests, ANOVA and correlation depend on this
    assumption, we test each variable separately.
  >> VARIABLE SELECTION:
    - Variables: sprint_100m_sec
    - Variables: VO2_max
    - Variables: lactate_mmol

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    3 continuous variables tested:
    sprint_100m_sec (100m time)  : Shapiro-Wilk = 0.992  p = 0.392   KS p = 0.393   Normal
    VO2_max                      : Shapiro-Wilk = 0.996  p = 0.890   KS p = 0.892   Normal
    lactate_mmol (lactate)       : Shapiro-Wilk = 0.855  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three continuous variables -- 100m sprint time, VO2max and blood lactate -- are normally
    distributed. The result splits instructively: sprint time and VO2max are normal (p > 0.39), but lactate deviates
    significantly from normality (p < .001). This is typical in physiological data -- lactate is often right-skewed
    (most athletes low, a few very high). Practical upshot: we can safely use parametric tests (t-test, ANOVA, Pearson)
    on sprint and VO2max; for lactate, nonparametric methods (Mann-Whitney/Kruskal-Wallis) or a transform are more
    appropriate. The normality check is a critical preliminary step that decides, per variable, which test family fits.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_sprint.xlsx
  >> SCENARIO (narration):
    We investigate whether the athletes' mean 100m sprint time differs from a 12.0 s
    reference threshold. With one group and a fixed reference, the one-sample t-test
    is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: sprint_100m_sec
    - Test value (mu): 12.0

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = -9.877   p < .001 ***   Cohen d = -0.988 (Large)
    Mean 100m = 11.38 s   (test mu = 12.0, reference threshold)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the athletes' mean 100m sprint time against a reference threshold of 12.0 seconds. The result is very
    strong: mean 11.38 s, clearly below the threshold (i.e. faster) -- t(99) = -9.88, p < .001, d = -0.99, a large
    effect. These athletes are too fast to be chance relative to the reference. The one-sample t-test is the right way
    to compare a performance measure against a known standard/norm (reference time, target mark); in sports science it
    is widely used to evaluate a performance against a benchmark.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent-Samples t-Test
    file: 04_independent_t_amateur_pro.xlsx
  >> SCENARIO (narration):
    We compare the mean VO2max of two level groups (amateur/professional). With two
    separate groups and a continuous measure, the independent-samples t-test is
    appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): level
    - Test variable: VO2_max

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = -7.103   p < .001 ***   Cohen d = -1.326 (Large)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the mean VO2max of two level groups (amateur/professional). The difference is significant and large:
    t(113) = -7.10, p < .001, d = -1.33. It shows the between-group difference is too pronounced to be chance and is
    also practically noteworthy. The independent-samples t-test is the standard way to compare the means of two
    separate groups (amateur/pro, two disciplines, two teams) on a continuous measure; it is a fundamental tool for
    detecting group differences in sports science.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired-Samples t-Test
    file: 05_paired_t_force.xlsx
  >> SCENARIO (narration):
    We compare force measured before and after a program in the same athletes. Since
    the measures are paired, the paired t-test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: force_before_kg
    - 2nd measure / group: force_post_kg

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(49) = -24.94   p < .001 ***   Cohen d_z = -3.526 (Large)
    Mean diff (before - after) = -15.58 kg   H0 REJECTED

>> COMMENTARY (narration):
    We paired and compared force (kg) measured before and after a training program in the same athletes. The result is
    very strong: mean difference -15.6 kg, t(49) = -24.94, p < .001, d_z = -3.53, a huge effect. Force rose
    significantly and substantially after the program. The paired t-test compares two timed measures on the same unit
    (before/after program); by isolating individual change it is more powerful than the independent test and is the
    right way to measure a training effect in sports science.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_discipline_VO2.xlsx
  >> SCENARIO (narration):
    We compare the effect of four disciplines on mean VO2max. With more than two
    groups, one-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: VO2_max
    - Factor (categorical): discipline

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,136) = 51.60   p < .001 ***   eta^2 = 0.532   H0 REJECTED

>> COMMENTARY (narration):
    We compared mean VO2max across four disciplines. The result is significant and very strong: F(3,136) = 51.60,
    p < .001, eta^2 = 0.53 -- half of VO2max variance comes from discipline differences. At least one discipline
    differs significantly. One-way ANOVA compares the means of more than two groups at once (avoiding the error
    inflation of many t-tests); in sports science it is the core method for comparing different discipline/level/group
    performance. Which pairs differ is then determined by post-hoc tests.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova_discipline_sex.xlsx
  >> SCENARIO (narration):
    We examine the main effects and interaction of discipline and sex on performance
    simultaneously. With two categorical factors, two-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Factor (categorical): discipline
    - 2nd Factor: sex

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    discipline (main effect): F(2) = 21.43   p < .001 ***   eta^2p = 0.273
    sex (main effect)       : F(1) = 36.44   p < .001 ***   eta^2p = 0.242

>> COMMENTARY (narration):
    We examined two factors at once: how do discipline and sex affect performance? Both main effects are significant
    and strong (discipline: F(2) = 21.43, eta^2p = 0.27; sex: F(1) = 36.44, eta^2p = 0.24). Both discipline and sex
    make an independent contribution to performance. The power of two-way ANOVA is that it tests both factors and their
    interaction in a single model -- isolating each factor's pure effect with the other controlled. It is ideal for
    answering "which variable really makes a difference?" in sports science.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated-Measures ANOVA
    file: 08_repeated_anova_training.xlsx
  >> SCENARIO (narration):
    We compare performance over four consecutive training periods in the same
    athletes. With repeated measures on the same unit, repeated-measures ANOVA is
    appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: measurement_1
    - Repeated measures: measurement_2
    - Repeated measures: measurement_3
    - Repeated measures: measurement_4

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,177) = 108.40   p < .001 ***   eta^2p = 0.648   n = 60   H0 REJECTED

>> COMMENTARY (narration):
    We compared performance measured across four consecutive training periods in the same 60 athletes. The result is
    very strong: F(3,177) = 108.40, p < .001, eta^2p = 0.65 -- the between-period difference is huge and most of the
    effect is time-related. Athlete performance changes significantly across periods. Repeated-measures ANOVA compares
    three or more timed measurements on the same unit; by holding individual differences constant it yields high
    statistical power and is the right choice for multi-period development tracking (training cycle, periodic testing)
    in sports science.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_discipline_3DV.xlsx
  >> SCENARIO (narration):
    We test discipline's effect on sprint, running and force at once. With several
    correlated dependent variables, MANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: sprint_sec
    - Dependent variable: running_min
    - Dependent variable: force_kg
    - Factor (categorical): discipline

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.065   F(6,230) = 111.95   p < .001 ***   n = 120
    Dependent: sprint_sec, running_min, force_kg   |   Factor: discipline   H0 REJECTED

>> COMMENTARY (narration):
    We tested discipline's effect on three dependent variables (sprint time, running, force) simultaneously. With
    Wilks' Lambda = 0.07, F(6,230) = 111.95, p < .001, the effect is very strong. MANOVA examines several correlated
    outcomes in one test, both preventing the error inflation of many separate ANOVAs and capturing the joint
    information the variables carry together. In sports science, when discipline affects not a single measure but the
    performance "bundle" (sprint-running-force) together, MANOVA reveals this multivariate difference as a single
    decision.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova_program_VO2.xlsx
  >> SCENARIO (narration):
    We compare post-program VO2max across groups while controlling baseline VO2max
    as a covariate. With a confounding continuous variable, ANCOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: VO2_after
    - Factor (categorical): program
    - Covariate: VO2_baseline

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group effect significant   eta^2p = 0.680 (very large)   covariate: VO2_baseline   H0 REJECTED

>> COMMENTARY (narration):
    We compared post-program VO2max (VO2_after) across groups while controlling baseline VO2max (VO2_baseline) as a
    covariate. This rules out the objection that "the groups differed at baseline" and measures the pure group effect:
    the effect is very large (eta^2p = 0.68). ANCOVA makes the group comparison fair by statistically holding a
    confounding continuous variable constant -- it is the answer to "once we equalize baseline conditioning, does the
    program still make a difference?" and is the standard tool in baseline/post-measure designs in sports science.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci_lactate.xlsx
  >> SCENARIO (narration):
    For mean blood lactate we build a confidence interval via resampling, with no
    distributional assumption. For skewed data, bootstrap is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: lactate_mmol

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean (lactate, mmol) = 5.02
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    For the mean blood lactate (mmol) we produced a 95% confidence interval via resampling -- with no distributional
    assumption: observed mean 5.02 mmol. Bootstrap builds the sampling distribution of the statistic empirically by
    resampling the data thousands of times from itself; it is a reliable way to give a confidence interval when
    normality does not hold (lactate is skewed) or no formula is known. In sports science it provides more robust
    estimates than classic t-intervals for skewed physiological measures.

====================================================================================

#12  Permutation Test
    file: 12_permutation_method_injury.xlsx
  >> SCENARIO (narration):
    We test the injury-ratio difference between two training groups via permutation,
    with no distributional assumption. For small samples/odd distributions,
    permutation is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: injury_ratio
    - Grouping (categorical): training

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = 8.037   p < .001 ***   significant mean difference between the two groups   H0 REJECTED

>> COMMENTARY (narration):
    We tested the injury-ratio difference between two training groups without any distributional assumption, via
    permutation: the null distribution was built by randomly swapping group labels, and the observed difference turned
    out very rare in that distribution (p < .001). Because the permutation test is exact and distribution-free, it is a
    safe alternative to the parametric t-test for small samples or odd distributions; in sports science it is a robust
    choice for group comparisons where assumptions are in doubt.

====================================================================================

#13  Multiple Comparison
    file: 13_multiple_comparison_program.xlsx
  >> SCENARIO (narration):
    We take the p-values of eight training-program comparisons together and apply
    multiple-testing correction. With many tests, p-adjustment is appropriate.
  >> VARIABLE SELECTION:
    - Variables: program
    - Variables: performance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Raw: 6/8 significant   |   After Bonferroni / Holm / BH (FDR): 6/8 significant (all retained)

>> COMMENTARY (narration):
    We took the p-values of eight separate training-program comparisons together and applied multiple-testing
    correction. Raw, 6 comparisons were significant; after Bonferroni/Holm/BH all 6 remained significant -- so these
    differences are strong enough to survive correction. When many tests are run, the rate of false positives that look
    "significant" by chance alone inflates; correction methods tighten the threshold to control this error. In sports
    science, when many programs/methods are compared at once, this step protects inference reliability.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney_level_satisfaction.xlsx
  >> SCENARIO (narration):
    We compare the satisfaction distribution of two level groups. Since normality
    fails, Mann-Whitney is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): level
    - Test variable: satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    U = 382.00   p < .001 ***   r = 0.58   H0 REJECTED

>> COMMENTARY (narration):
    We compared the satisfaction distribution of two level groups (level) based on ranks rather than means: U = 382,
    p < .001, large effect (r = 0.58). The difference is very pronounced. Mann-Whitney is the nonparametric counterpart
    of the t-test; when normality fails or the scale is ordinal, it is the right way to compare two groups. In sports
    science it is a robust choice for ordinal measures like satisfaction/grade.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank
    file: 15_wilcoxon_psychology.xlsx
  >> SCENARIO (narration):
    We compare attitude scores measured before/after in the same athletes, without
    assuming normality. For paired non-normal data, Wilcoxon is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: attitude_before
    - 2nd measure / group: attitude_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    W = 0.00   p < .001 ***   r = 0.89   n (non-zero) = 34   H0 REJECTED

>> COMMENTARY (narration):
    We compared attitude scores measured before and after on the same athletes (before/after a psychology course) --
    without assuming normality, on a rank basis: W = 0, p < .001, very large effect (r = 0.89). Attitude changed
    consistently and strongly after the intervention. Wilcoxon is the nonparametric counterpart of the paired t-test;
    it is the right choice for ordinal or non-normal before/after measures. In sports science it gives reliable results
    for pre/post psychological-intervention measures.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_level_flexibility.xlsx
  >> SCENARIO (narration):
    We compare the flexibility distribution across three level groups. Since
    normality/variance homogeneity fails, Kruskal-Wallis is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): level
    - Test variable: flexibility_cm

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 9.16   p = 0.010 *   eta^2_H = 0.17   H0 REJECTED

>> COMMENTARY (narration):
    We compared the flexibility (flexibility_cm) distribution across three level groups by ranks rather than means:
    H(2) = 9.16, p = 0.010, eta^2_H = 0.17, a moderate effect. At least one group differs significantly. Kruskal-Wallis
    is the nonparametric counterpart of one-way ANOVA; when normality or variance homogeneity fails, it is the right way
    to do multi-group comparison. In sports science it is robust for detecting group differences on skewed or ordinal
    measures.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman_jury_teknik.xlsx
  >> SCENARIO (narration):
    We compare jury scores measured under four conditions/techniques in the same
    units. As a nonparametric repeated measure, Friedman is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: condition_A
    - Repeated measures: condition_B
    - Repeated measures: condition_C
    - Repeated measures: condition_D

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 39.08   p < .001 ***   Kendall W = 0.434   n = 30   H0 REJECTED

>> COMMENTARY (narration):
    We compared jury scores measured under four conditions/techniques (condition_A..D) in the same 30 units -- as a
    nonparametric repeated measure: chi^2(3) = 39.08, p < .001, Kendall W = 0.43 (moderate concordance). The
    between-condition difference is significant. Friedman is the nonparametric counterpart of repeated-measures ANOVA;
    it is the right choice for ordinal or non-normal paired multi-condition measures. In sports science it is used to
    compare the same athlete's rankings across different techniques/conditions.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_galibiyet.xlsx
  >> SCENARIO (narration):
    We test the proportion of matches ending in a win against an expected 50%. With
    a binary outcome and a theoretical proportion, the binomial test is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: winner_geldi
    - Expected proportion: 0.50
    - Success value: 1

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed proportion = 0.714   (expected p = 0.50)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested the proportion of matches ending in a win against an expected 50%: observed proportion 71.4%, well above
    expectation (p < .001). The win rate is too high to be chance. The binomial test is the exact method for comparing
    the observed proportion of a binary (win/loss) outcome with a theoretical proportion; in sports science it directly
    tests whether win/success rates meet a target or a 50:50 expectation.

====================================================================================

#19  Sign Test
    file: 19_sign_test_heart_atisi.xlsx
  >> SCENARIO (narration):
    We look at the direction of before/after heart rate in the same athletes. When
    only directional information is reliable, the sign test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: heart_before_bpm
    - 2nd measure / group: heart_post_bpm

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive diffs (before > after) = 25   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We looked at the direction of change in heart rate measured before/after on the same athletes (heart_before vs
    post bpm): the large majority of changes go in one direction, p < .001 for a significant directional change. The
    sign test uses only the direction of the difference (increase/decrease), not its magnitude, so it is the
    before/after test that requires the fewest assumptions. In sports science it is a robust choice when the measure is
    skewed or only directional information is reliable (did the heart rate drop or rise).

====================================================================================

#20  Runs Test
    file: 20_runs_test_season_score.xlsx
  >> SCENARIO (narration):
    We test whether the season score sequence (around the median) is random. For
    sequence randomness, the runs test is appropriate.
  >> VARIABLE SELECTION:
    - Column: score_avg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 28   Z = -0.781   p = 0.435 ns   data may be considered random

>> COMMENTARY (narration):
    We tested whether the season-long score sequence (score_avg, around the median) is random: 28 runs, Z = -0.78,
    p = 0.44 -- no pattern, the sequence is random. The runs test checks whether values in a sequence form a systematic
    pattern (clusters, cycles, win/loss streaks). In sports science it is used to determine whether systematic patterns
    (form streaks) or randomness dominate in season-performance series, and in checking the independence assumption.

====================================================================================

#21  Chi-Square Independence
    file: 21_chisquare_discipline_injury.xlsx
  >> SCENARIO (narration):
    We test whether discipline and injury status are related. With two categorical
    variables, chi-square is appropriate.
  >> VARIABLE SELECTION:
    - Variables: discipline
    - Variables: status

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(2) = 2.42   p = 0.298 ns   Cramer's V = 0.056 (Weak)   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether two categorical variables -- discipline and injury status (status) -- are related: chi^2(2) =
    2.42, p = 0.30, Cramer's V = 0.06 -- no significant relationship, discipline and injury are largely independent.
    The chi-square test of independence detects the relationship between two qualitative variables from a cross-tab; in
    sports science it tests categorical relationships such as discipline-injury, group-outcome. Here "no relationship"
    is also a valuable finding (injury risk does not vary by discipline).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_position.xlsx
  >> SCENARIO (narration):
    We test whether a four-category position's observed distribution fits an equal
    expected distribution. For one categorical variable, goodness-of-fit is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: position

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 27.56   p < .001 ***   N = 200, k = 4 (expected: equal distribution)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the observed distribution of a four-category variable (position) fits an equal (1/k) expected
    distribution: chi^2(3) = 27.56, p < .001 -- the positions are not equally distributed, some are clearly more
    frequent. The goodness-of-fit test compares the observed frequencies of a single categorical variable with a
    theoretical expectation (equal proportions, a known ratio); in sports science it tests whether position/discipline
    distributions match an expected profile.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_method_recovery.xlsx
  >> SCENARIO (narration):
    In a 2x2 table we test the method-recovery relationship exactly, due to small
    cell frequencies. For few observations, Fisher is appropriate.
  >> VARIABLE SELECTION:
    - Variables: method
    - Variables: recovery

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher exact p = 0.151 ns   Odds Ratio = 3.375   Phi = 0.292 (Moderate)   H0 NOT REJECTED

>> COMMENTARY (narration):
    In a 2x2 cross-tab (method x recovery) we tested the relationship exactly -- using Fisher instead of chi-square
    because of small cell frequencies: p = 0.151, no significant relationship (Phi = 0.29). Fisher's exact test is the
    right choice when expected frequencies are low and the chi-square approximation is unreliable; it computes the
    probability exactly rather than approximately. In sports science it gives reliable results for small-sample pilot
    comparisons (does the new recovery method work).

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_test_comparison.xlsx
  >> SCENARIO (narration):
    We compare the binary outcome of two tests paired in the same athletes. For
    paired binary change, McNemar is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: test_a
    - 2nd measure / group: test_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar (exact) min(b,c) = 3   p = 0.092 ns   discordant b = 3, c = 10   OR = 0.30   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested the direction of change in a binary outcome of two tests (test_a vs test_b) paired in the same athletes:
    although the changes are imbalanced (b = 3, c = 10), p = 0.092 is non-significant. McNemar tests the before/after
    (or two-test) binary status change in the same unit and looks only at discordant pairs. In sports science it is the
    right tool for detecting whether two tests'/methods' binary outcomes differ systematically; the exact (binomial)
    version is used for small discordant counts.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_jury_goodnessfit.xlsx
  >> SCENARIO (narration):
    We measure the agreement of two juries classifying the same athletes. For
    categorical agreement, kappa is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: jury_a
    - 2nd measure / group: jury_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Kappa = 0.736 (substantial agreement)

>> COMMENTARY (narration):
    We measured the agreement of two juries (jury_a vs jury_b) classifying the same athletes into the same categories
    -- excluding chance agreement: kappa = 0.74, substantial. Unlike raw percent agreement, kappa reports categorical
    agreement after removing the chance-agreement share, so it is more honest. In sports science it is the standard
    index for measuring how consistent two juries'/referees' classifications (technical score, grade) are.

====================================================================================

#26  Cochran-Mantel-Haenszel
    file: 26_cmh_age_group_recovery.xlsx
  >> SCENARIO (narration):
    We test the group-recovery relationship controlling for age-group strata. For a
    stratum-controlled relationship, CMH is appropriate.
  >> VARIABLE SELECTION:
    - Variables: group
    - Variables: recovery
    - Stratum: age_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi^2(1) = 49.34   p < .001 ***   Common OR (MH) = 5.04 [3.16, 8.03]
    Breslow-Day p = 0.382 (OR homogeneous)   H0 REJECTED

>> COMMENTARY (narration):
    We tested the group-recovery relationship while controlling for age-group strata: the common Odds Ratio across
    strata = 5.04 [3.16, 8.03], p < .001; Breslow-Day p = 0.38 means this relationship is consistent across all strata.
    CMH measures the pure strength of the association by holding a confounding stratum variable constant (preventing
    Simpson's paradox). In sports science it is ideal for robustly estimating a group-outcome (method-recovery)
    relationship while controlling age differences.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_sex_discipline_frequency.xlsx
  >> SCENARIO (narration):
    We model the joint relationship structure of three categorical variables. For
    more than two categorical dimensions, log-linear is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sex
    - Variables: discipline_type
    - Variables: training_frequency

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 50.54   Pearson chi^2 = 0.456   joint relationship of 3 categorical variables (sex, discipline_type, training_frequency) modeled

>> COMMENTARY (narration):
    We examined the joint relationship structure of three categorical variables (sex, discipline type, training
    frequency) with a log-linear model. The model explains cell frequencies via main effects and interactions; it
    reveals which pairs/triples of variables vary together. It is the multivariable generalization of the two-way
    cross-tab. In sports science it is used to analyze the joint dependency pattern of more than three categorical
    dimensions (sex x discipline x frequency), balancing parsimony with AIC to select the most explanatory structure.

====================================================================================

#28  Cross-Tabulation
    file: 28_cross_age_motivation.xlsx
  >> SCENARIO (narration):
    We examine age group and motivation in a cross-tab. For the relationship of two
    qualitative variables, cross-tab is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age_group
    - Variables: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(6) = 27.24   p < .001 ***   Cramer's V = 0.197 (Moderate)   H0 REJECTED

>> COMMENTARY (narration):
    We examined two categorical variables (age group x motivation) in a cross-tab: chi^2(6) = 27.24, p < .001, V = 0.20
    -- a significant, moderate relationship. Cross-tab + chi-square shows the direction and strength of the association
    between two qualitative variables at the cell level. In sports science it is used to describe age-motivation,
    demography-attitude relationships; Cramer's V (0.20) measures the practical strength.

====================================================================================

#29  Multiple Response - Frequency
    file: 29_mr_frequency_sports.xlsx
  >> SCENARIO (narration):
    We analyze a multi-select sports question. For a multi-select question,
    multiple-response frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sports (coklu yanit / multi-response)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)   multi-select "sports" question

>> COMMENTARY (narration):
    We analyzed a question where several options could be checked ("which sports are you interested in?"): each of 250
    participants gave one or more sports. Multiple-response frequency analysis gives the count of checks per option and
    both the response and case percentages separately (percentages sum to over 100, because one person picks multiple).
    In sports science it is the standard method for correctly summarizing multi-select questions (sports practiced,
    equipment used).

====================================================================================

#30  Multiple Response - Cross-Tab
    file: 30_mr_categorical_sports_sex.xlsx
  >> SCENARIO (narration):
    We cross-tabulate the multi-response sports question by sex. To break a
    multi-select by a category, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sports (coklu yanit / multi-response)
    - Grouping (categorical): sex

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: sex   |   multi-response "sports" x sex

>> COMMENTARY (narration):
    We cross-tabulated the multi-response sports question by a single category (sex): we compared the per-sport check
    rates for each group. A multiple-response cross-tab answers "does one group check certain options more often than
    another?". In sports science it is used to compare groups' (sex, age) multi-select preference profiles (sports
    practiced, interests).

====================================================================================

#31  Multiple Response x Multiple Response
    file: 31_mr_mr_sport_equipment.xlsx
  >> SCENARIO (narration):
    We cross-tabulate two multi-response questions (sports x equipment) against each
    other. For many-to-many co-occurrence, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sports (coklu / multi)
    - Variables: equipment (coklu / multi)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180   two multi-responses (sports x equipment) cross-tabulated

>> COMMENTARY (narration):
    We cross-tabulated two separate multi-response questions (sports x equipment) against each other. This is the most
    complex form of tabulation: on both axes a unit contributes to more than one cell. Multiple-response by
    multiple-response reveals many-to-many co-occurrences such as "which sports appear together with which equipment?".
    In sports science it is used to examine the matching pattern of nested preference bundles (sport practiced x
    equipment used).

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q_weather_performance.xlsx
  >> SCENARIO (narration):
    We test whether a binary outcome varies across weather conditions in the same
    athletes. For 3+ repeated binary measures, Cochran's Q is appropriate.
  >> VARIABLE SELECTION:
    - Columns: hava-bazli ikili sutunlar / weather-wise binary columns

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000 ns   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the binary outcome (good performance yes/no) measured under different weather conditions in the
    same athletes varies significantly from condition to condition: Q = 0, p = 1.00 -- no difference across weather
    conditions. Cochran's Q is the generalization of McNemar to more than two repeated conditions; it compares 3+ binary
    measures in the same unit. In sports science it is the right method for comparing whether the same athletes perform
    well across different conditions (weather, surface).

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5degisken.xlsx
  >> SCENARIO (narration):
    We examine all pairwise correlations among five continuous variables. For
    relationship direction and strength, the correlation matrix is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: training_hour
    - Variables: VO2_max
    - Variables: sprint_100m_sec
    - Variables: fat_ratio_pct

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables: age, training_hour, VO2_max, sprint_100m_sec, fat_ratio_pct
    Strong relationship (|r| >= 0.75): NONE above this threshold

>> COMMENTARY (narration):
    We computed all pairwise Pearson correlations among five continuous variables (age, training hours, VO2max, 100m
    time, fat ratio): none exceeded the 0.75 strong threshold -- the variables move largely independently. This too is
    a valuable finding: most measures carry different information, none substitutes for another. The correlation matrix
    summarizes the direction and strength of relationships at a glance; in sports science it is the first step in
    spotting overlap among indicators (multicollinearity risk) and seeing which variables truly give separate signals.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Agreement
    file: 34_bland_altman_VO2_method.xlsx
  >> SCENARIO (narration):
    We examine the agreement of two methods (pulse oximetry/gas analysis) measuring
    the same VO2. For method interchangeability, Bland-Altman is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: VO2_pulse_oksimetre
    - 2nd measure / group: VO2_gas_analysis

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pulse oximetry (VO2_pulse): mean = 73.96   |   Gas analysis (VO2_gas): mean = 71.49
    Bias (mean difference) ~ 2.48   |   limits of agreement (LoA) computed

>> COMMENTARY (narration):
    We examined how well two methods measuring the same VO2 (pulse oximetry vs gas analysis) agree: the mean systematic
    difference (bias) is ~2.5 units, and 95% limits of agreement were reported. Unlike correlation, Bland-Altman
    answers "can the two methods be used interchangeably?" -- high correlation does not mean agreement, there may be a
    systematic shift. In sports science it is the standard method for testing the interchangeability of two measuring
    methods/devices.

====================================================================================

#35  Effect Size
    file: 35_effect_size_program.xlsx
  >> SCENARIO (narration):
    We measure the practical size of the performance-increase difference between two
    programs. For importance beyond the p-value, effect size is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: performance_increase
    - Grouping (categorical): program

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = 1.442 (very large effect)   performance-increase difference between two programs (program)

>> COMMENTARY (narration):
    We measured the practical size of the performance-increase difference between two training programs, independent of
    the p-value, with Cohen's d: d = 1.44, a very large effect. The p-value answers "is there a difference?"; effect
    size answers "how important is it?". Because large samples can produce significant but trivial differences,
    reporting effect size is essential. In sports science this measure clarifies whether the difference between two
    programs/methods is practically noteworthy.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_physical_psychology.xlsx
  >> SCENARIO (narration):
    We examine the joint structure between the physical-capacity set and the
    psychological set. For the relationship between two multivariate sets, CCA is
    appropriate.
  >> VARIABLE SELECTION:
    - X variables: VO2
    - X variables: force
    - X variables: flexibility
    - X variables: balance
    - X variables: reaction
    - Y variables: motivation
    - Y variables: concentration
    - Y variables: self_efficacy
    - Y variables: competition

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.980   chi^2(20) = 1186.45   p < .001 ***
    CC2: r = 0.960   chi^2(12) = 724.85    p < .001 ***
    Set-X (physical): VO2, force, flexibility, balance, reaction  |  Set-Y (psychological): motivation, concentration, self_efficacy, competition

>> COMMENTARY (narration):
    We resolved the joint structure between two multivariate measure sets -- physical capacity (VO2, force, flexibility,
    balance, reaction) and psychological variables (motivation, concentration, self-efficacy, competition) -- with
    canonical correlation. The first two canonical functions are very strong (r = 0.980 and 0.960, p < .001): the two
    sets are intensely related. CCA answers "how is one variable set related to another?" in a single step -- it is the
    multivariate-on-both-sides version of multiple regression. In sports science it reveals the latent relationship
    structure between the physical bundle and the psychological bundle.

====================================================================================

#37  Correspondence Analysis
    file: 37_ca_discipline_time.xlsx
  >> SCENARIO (narration):
    We map the relationship between discipline and training time. To see the
    structure of two categorical variables, CA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: discipline
    - Variables: training_time

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.043   (2 dimensions)   discipline x training_time

>> COMMENTARY (narration):
    We mapped the relationship in the cross-tab of two categorical variables (discipline x training time) into a visual
    space with correspondence analysis: total inertia 0.043 (weak relationship) resolved over two dimensions. CA
    positions the chi-square relationship in a two-dimensional space, showing which categories are close (co-occurring).
    In sports science it is powerful for visually interpreting the structure of qualitative relationships such as
    discipline-time, discipline-preference; the low inertia indicates the relationship is weak.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_18madde.xlsx
  >> SCENARIO (narration):
    We cluster eighteen items (physical/psychological/tactical) by their similarity.
    To find the latent dimension structure, VarClus is appropriate.
  >> VARIABLE SELECTION:
    - Variables: physical_m1..6, psychological_m1..6, tactic_m1..6 (18 madde / items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: physical_m1..6, psychological_m1..6, tactic_m1..6)

>> COMMENTARY (narration):
    We clustered eighteen measurement items (physical, psychological, tactical, 6 each) by how related they are: they
    grouped into 3 main dimensions -- most likely matching the natural physical/psychological/tactical groups. VarClus
    groups the VARIABLES, not the observations -- by placing highly correlated items in the same cluster it reveals the
    latent dimensional structure of the data set. In sports science it is practical for reducing a long indicator set to
    a few core dimensions and for spotting redundant items.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_performance.xlsx
  >> SCENARIO (narration):
    We model performance with four predictors (age, training hours, motivation,
    BMI). To explain a continuous outcome with multiple variables, multiple
    regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): age
    - Predictor(s): training_hour
    - Predictor(s): motivation
    - Predictor(s): BMI

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.860   Adj. R^2 = 0.857   predictors: age, training_hour, motivation, BMI

>> COMMENTARY (narration):
    We modeled performance with four predictors (age, training hours, motivation, BMI) at once: the model explains 86%
    of variance (Adj. R^2 = 0.86) -- very strong explanatory power. Multiple regression gives each predictor's pure
    contribution to performance with the others held constant; thus it answers "which factor really raises
    performance?" while controlling confounders. In sports science it is the core method for identifying the drivers of
    performance and for prediction.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_injury.xlsx
  >> SCENARIO (narration):
    We model new injury (yes/no) with three predictors. For a binary outcome,
    logistic regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: new_injury
    - Predictor(s): age
    - Predictor(s): training_hour
    - Predictor(s): past_injury

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 = 0.074   binary outcome: new_injury   predictors: age, training_hour, past_injury

>> COMMENTARY (narration):
    We modeled a binary outcome (new injury yes/no) with three predictors: the model has weak explanatory power (pseudo
    R^2 = 0.07) and gives each predictor's effect on the odds. Logistic regression replaces linear regression when the
    outcome is binary; coefficients are converted to Odds Ratios to read "how many times does the injury odds change
    per unit increase in this variable?". In sports science it is the core model for predicting yes/no outcomes such as
    injury, success; here the explanatory power is low, suggesting additional predictors may be needed.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Count (Poisson) Regression
    file: 41_poisson_injury_count.xlsx
  >> SCENARIO (narration):
    We model the annual injury count with age and training hours. For a count
    outcome, Poisson regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: annual_injury_count
    - Predictor(s): age
    - Predictor(s): training_hour

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 491.28   Deviance = 212.93   count outcome: annual_injury_count   predictors: age, training_hour

>> COMMENTARY (narration):
    We modeled a count variable (annual injury count) with age and training hours. Poisson regression is the right
    model when the outcome is a count (0,1,2,... items); linear regression is unsuitable because it can produce
    negative/fractional predictions. Coefficients give the effect on the count rate. In sports science it is used to
    explain count outcomes such as injury count, match count, event tallies.

====================================================================================

#42  Multinomial Logistic
    file: 42_multinomial_discipline.xlsx
  >> SCENARIO (narration):
    We model the multi-category selected discipline with two predictors. For a
    nominal multi-class outcome, multinomial logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: selected_discipline
    - Predictor(s): height_cm
    - Predictor(s): force_kg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 472.13   Reference class: 'gymnastics'   outcome: selected_discipline (>2 categories)   predictors: height_cm, force_kg

>> COMMENTARY (narration):
    We modeled a more-than-two-category outcome (selected discipline) with two predictors (height, force). Multinomial
    logistic compares each category against a reference class (here 'gymnastics') with a separate logistic equation;
    coefficients are read as "as X increases, how does the chance of being in this discipline change relative to the
    reference?". When the outcome is nominal with more than two classes (discipline, position preference) it is the
    right choice. In sports science it is the standard model for multi-option class prediction (talent-discipline
    matching).

====================================================================================

#43  Ordinal Logistic
    file: 43_ordinal_performance.xlsx
  >> SCENARIO (narration):
    We model the ordinal performance level with training hours and motivation. For
    an ordinal outcome, ordinal logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance_level
    - Predictor(s): training_hour
    - Predictor(s): motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 402.22   ordinal outcome: performance_level   predictors: training_hour, motivation

>> COMMENTARY (narration):
    We modeled an ordinal outcome (performance level: low/medium/high) with training hours and motivation. Ordinal
    logistic uses the ORDER information between categories (which multinomial ignores); with a "proportional odds"
    assumption it explains all thresholds with one coefficient set. In sports science it is the right and more powerful
    choice for modeling naturally ordered outcomes such as performance level, grade, class.

====================================================================================

#44  PLS Regression
    file: 44_pls_sensor_performance.xlsx
  >> SCENARIO (narration):
    We predict performance from 12 wearable-sensor readings. For many highly
    correlated predictors, PLS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): sensor_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 (train) = 0.878   R^2 (5-fold CV) = 0.870   predictors: sensor_01..12 (12 sensors)

>> COMMENTARY (narration):
    We predicted performance from 12 wearable-sensor readings. Because the sensor readings are highly correlated
    (multicollinearity), classic regression becomes unstable; PLS reduces them to a few latent components and regresses
    on those. With cross-validated R^2 = 0.87, the model is both strong and generalizable. PLS is ideal when predictors
    are numerous or highly correlated; in sports science it is widely used for predicting performance from multi-sensor
    wearable data.

====================================================================================

#45  Probit Regression
    file: 45_probit_creatine.xlsx
  >> SCENARIO (narration):
    We estimate the probability of a performance increase from daily creatine dose
    with a probit model. For a binary dose-response outcome, probit is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance_increase_present
    - Predictor(s): creatine_day_g

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 161.12   Pseudo R^2 (McFadden) = 0.367   binary outcome: performance_increase_present   predictor: creatine_day_g

>> COMMENTARY (narration):
    We estimated the dose-response relationship -- the probability of a performance increase from daily creatine dose
    (creatine_day_g) -- with a probit model: the model is strong (pseudo R^2 = 0.37). Probit, like logistic, applies to
    binary outcomes; the difference is that its link function is the normal distribution. Probit is classic in
    dose-response studies. In sports science it is a robust choice for modeling the relationship between an ergogenic
    aid/dose and a binary response (increase yes/no).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_income.xlsx
  >> SCENARIO (narration):
    We model a threshold-piled annual-income variable with sponsor presence and
    level. For a censored outcome, tobit is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: annual_income_TL
    - Predictor(s): sponsor_present
    - Predictor(s): level

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 3959.69   Left censor   outcome: annual_income_TL (censored)   predictors: sponsor_present, level

>> COMMENTARY (narration):
    We modeled a lower-bounded (threshold-piled) annual-income variable with sponsor presence and level. Tobit is for
    "censored" dependent variables that pile up at a threshold; ordinary regression gives biased estimates by ignoring
    this pile-up. In sports science/sports economics it is the right model for floor-effect outcomes (zero/below-
    threshold income); coefficients reflect the true (uncensored) relationship.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_nutrition.xlsx
  >> SCENARIO (narration):
    We model performance with nutrition variables (protein, carbohydrate g/kg) in a
    Bayesian framework. To express uncertainty probabilistically, Bayesian
    regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): protein_g_kg
    - Predictor(s): karbon_g_kg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2 posterior mean = 26.40 (sd = 3.77)   outcome: performance   predictors: protein_g_kg, karbon_g_kg
    coefficient posteriors + P(beta>0) reported

>> COMMENTARY (narration):
    We modeled performance with nutrition variables (protein, carbohydrate g/kg) in a Bayesian framework: instead of
    point estimates we obtained each coefficient's full posterior distribution and the "probability the effect is
    positive". The Bayesian approach expresses uncertainty directly in probability language and can incorporate prior
    knowledge. In sports science, when the sample is small or prior-study information is valuable, it offers intuitive
    interpretations like "the nutrition effect is probably positive".

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear_age_performance.xlsx
  >> SCENARIO (narration):
    We model the S-shaped relationship of peak performance with age via a logistic
    growth curve. For a curved relationship, nonlinear regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance_peak
    - Predictor(s): age_year
    - Value: fonksiyon/function: logistic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Function: y = K / (1 + exp(-r*(x-x0)))  (logistic growth)   R^2 = 0.963

>> COMMENTARY (narration):
    We modeled the relationship of peak performance (performance_peak) with age not as a straight line but as an
    S-shaped logistic growth curve: the fit is very high (R^2 = 0.96). Nonlinear regression fits a theoretical function
    form directly to the data when the relationship is curved (development, saturation, threshold) and makes the
    parameters (ceiling K, rate r, inflection x0) interpretable. In sports science it is the right tool for modeling
    age-related performance/capacity development curves and development-saturation processes.

====================================================================================

#49  Ridge Regression
    file: 49_ridge_15olcum.xlsx
  >> SCENARIO (narration):
    We predict performance from 15 correlated measurements with ridge. For
    multicollinearity, ridge is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): measurement_01..15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 1.0   R^2 = 0.778   predictors: measurement_01..15 (15 measurements)

>> COMMENTARY (narration):
    We predicted performance from 15 correlated measurements with ridge regression (R^2 = 0.78). Ridge adds an L2
    penalty to shrink all coefficients in magnitude without zeroing them; this prevents the instability caused by high
    correlation among measurements (multicollinearity). It gives more stable and generalizable estimates where classic
    regression's coefficients balloon and flip sign. In sports science it is preferred for prediction with many
    co-varying physiological measurements.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso_30pred.xlsx
  >> SCENARIO (narration):
    We model the elite score from 30 candidate features with lasso, selecting the
    important ones. For automatic variable selection, lasso is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: elite_score
    - Predictor(s): x01..30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.1   R^2 = 0.719   19 of 30 features shrunk to zero (automatic variable selection)

>> COMMENTARY (narration):
    We modeled the elite score (elite_score) from 30 candidate features with lasso regression: R^2 = 0.72, and lasso
    shrank 19 of the 30 feature coefficients exactly to zero, selecting only the effective variables. This is its
    difference from ridge: because lasso can zero coefficients, it performs prediction and variable selection at the
    same time. In sports science it is very useful for automatically winnowing the "few truly important factors" out of
    many candidate indicators (a sparse, interpretable model).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_motivation_training.xlsx
  >> SCENARIO (narration):
    We test the motivation -> training minutes -> performance chain. To resolve the
    intermediate mechanism, mediation analysis is appropriate.
  >> VARIABLE SELECTION:
    - Predictor(s): motivation
    - Value: M (araci/mediator): training_minute
    - Dependent variable: performance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = -0.091   95% CI [-0.528, 0.000]   (borderline; not significant)
    X: motivation  M: training_minute  Y: performance

>> COMMENTARY (narration):
    We tested the chain "motivation (X) -> training minutes (M) -> performance (Y)": the indirect effect is -0.091, its
    95% CI includes zero (at the boundary) -- so motivation's indirect effect on performance through training time was
    not significant. Mediation analysis resolves "why/how does X affect Y?" through an intermediate mechanism; here this
    mechanistic path was not supported. In sports science it is powerful for understanding which intermediate process
    (training, effort) a factor's (motivation) effect flows through -- or does not.

====================================================================================

#52  Path Analysis
    file: 52_path_ability_performance.xlsx
  >> SCENARIO (narration):
    We test direct/indirect relationships as a single causal diagram. For a
    relationship network, path analysis is appropriate.
  >> VARIABLE SELECTION:
    - Value: performance ~ ability + motivation + effort
    - Value: motivation ~ ability

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.861   RMSEA = 0.271   model: performance ~ ability + motivation + effort; motivation ~ ability

>> COMMENTARY (narration):
    We tested the direct and indirect relationships among several variables as a single causal diagram. The fit indices
    (CFI = 0.86, RMSEA = 0.27) show the model fits the data partially, with room for improvement. Path analysis
    estimates the whole relationship network at once instead of separate regressions; it shows how variables affect
    each other and a common outcome. In sports science it is used to test theory-based relationship models (ability ->
    motivation -> performance).

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_athlete_term.xlsx
  >> SCENARIO (narration):
    We model performance measured repeatedly across periods in the same athletes,
    taking athlete as a random effect. For repeated/nested data, LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): term
    - Cluster: athlete_id (random)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC = 0.947   Group variance (random intercept) = 98.73   outcome: performance   fixed: term   group: athlete_id

>> COMMENTARY (narration):
    We modeled performance measured repeatedly across periods in the same athletes, taking athlete identity as a random
    effect. ICC = 0.95 is very high: almost all performance variability comes from between-athlete differences, while
    within-athlete periods are very similar. LMM correctly handles dependency in nested/repeated (measurements within
    athlete) data; it solves the "independence" assumption that ordinary regression violates via random effects. In
    sports science it is the right choice for panel/repeated-measure data.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation.xlsx
  >> SCENARIO (narration):
    Instead of deleting missing data we fill it with 5 plausible value sets. To
    handle missingness without bias, multiple imputation is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): age
    - Predictor(s): VO2
    - Predictor(s): force_kg

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (imputation) = 5   outcome: performance   predictors: age, VO2, force_kg

>> COMMENTARY (narration):
    Instead of deleting missing data, we analyzed by producing 5 plausible value sets (m = 5) and combining the
    results. Multiple imputation -- unlike filling gaps with a single estimate (which ignores uncertainty) -- accounts
    for imputation uncertainty too, yielding unbiased estimates and correct standard errors. In sports science it is
    the modern standard for handling measurement dropouts/missing test results without shrinking the sample or
    distorting results.

====================================================================================

#55  GEE
    file: 55_gee_program_performance.xlsx
  >> SCENARIO (narration):
    We model performance in repeated visit measures of the same athletes. For a
    population-average effect, GEE is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): visit
    - Cluster: athlete_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coef = 1.907   p < .001 ***   QIC = 305.61   outcome: performance   group: athlete_id

>> COMMENTARY (narration):
    We modeled performance in repeated visit measures of the same athletes; the visit effect is significant (b = 1.91,
    p < .001). GEE estimates the POPULATION-AVERAGE effect rather than individual effects in repeated/clustered data and
    corrects within-group correlation with a "working correlation structure". While LMM focuses on individual random
    effects, GEE focuses on the average trend. In sports science it is preferred for population-level questions like
    "what is the time/program effect in the average athlete?".

====================================================================================

#56  GLMM
    file: 56_glmm_rehabilitation.xlsx
  >> SCENARIO (narration):
    We model a repeatedly measured pain count in the same athletes with month and
    rehabilitation, taking athlete as a random effect. For repeated counts, GLMM is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: pain_count
    - Predictor(s): month
    - Predictor(s): rehabilitation
    - Cluster: athlete_id (Poisson)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rehabilitation coef = -0.423   p < .001 ***   outcome: pain_count (count, Poisson)   group: athlete_id

>> COMMENTARY (narration):
    We modeled a COUNT outcome (pain-event count) measured repeatedly in the same athletes with month and
    rehabilitation, taking athlete as a random effect; the rehabilitation effect is significant and negative
    (b = -0.42, p < .001) -- rehabilitation reduces pain count. GLMM extends LMM to non-normal outcomes (count, binary):
    it handles both the distribution (Poisson) and the clustering (random effect) at the same time. In sports science it
    is the right model for repeatedly measured count outcomes (monthly pain/injury count).

====================================================================================

#57  Elastic Net
    file: 57_elasticnet.xlsx
  >> SCENARIO (narration):
    We model a performance index from 40 predictors with Elastic Net. For many
    clustered predictors, Elastic Net is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance_index
    - Predictor(s): x01..40

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.114   L1 ratio = 0.50   R^2 (train) = 0.400   40 predictors

>> COMMENTARY (narration):
    We modeled a performance index (performance_index) from 40 predictors with Elastic Net. Elastic Net blends the
    ridge (L2) and lasso (L1) penalties (L1 ratio = 0.5): it both keeps groups of correlated variables together (ridge
    property) and zeroes out redundant ones (lasso property). It thus provides a balanced model with high-dimensional,
    clustered predictors. In sports science it is chosen when there are many clustered measurements, where lasso or
    ridge alone is insufficient.

====================================================================================

#58  Robust Regression
    file: 58_robust_area_yield.xlsx
  >> SCENARIO (narration):
    We model performance with age, down-weighting outliers. For data with outliers,
    robust regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust Intercept = 100.08   OLS Intercept = 96.57   outcome: performance   predictor: age

>> COMMENTARY (narration):
    When modeling performance with age, we used robust regression to prevent outliers from distorting the estimate. The
    gap between the robust and OLS intercepts (100.08 vs 96.57) shows a few outlying observations pull the classic
    estimate; the robust method down-weights them to reflect the "typical" relationship. In sports science it is the
    right way to get robust estimates without deleting outliers (abnormal test, erroneous record) in data that contains
    them.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_yield.xlsx
  >> SCENARIO (narration):
    We model different points of performance's distribution (lower 10%, median,
    upper 90%) separately. For varying effects, quantile regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): training_hour
    - Predictor(s): motivation
    - Quantiles: 0.10 / 0.50 / 0.90

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    separate slopes reported for q=0.10, q=0.50, q=0.90   outcome: performance   predictors: training_hour, motivation

>> COMMENTARY (narration):
    We modeled not just performance's mean but different points of the distribution (lower 10%, median, upper 90%)
    separately. Predictor effects can vary by quantile -- a factor may be strong for low-performing athletes and weak
    for high-performing ones. Quantile regression gives the true picture when the "mean effect" is misleading (effect
    varies across the distribution). In sports science it shows what classic regression misses by examining elite/
    amateur extreme-group behavior and performance variability.

====================================================================================

#60  ROC Curve
    file: 60_roc_quality.xlsx
  >> SCENARIO (narration):
    We assess how well a combined score separates a binary outcome with ROC. For
    discrimination and threshold selection, ROC is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: elite_label
    - Predictor(s): combined_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.969 (excellent discrimination)   Youden optimum threshold = 0.379   Sensitivity = 0.955, 1-Specificity = 0.108
    n = 250 (positive 111, negative 139)

>> COMMENTARY (narration):
    We assessed how well a combined score (combined_score) separates a binary outcome (elite_label) with a ROC curve:
    AUC = 0.97, excellent discrimination; the optimum decision threshold was set at 0.38 via Youden. ROC shows the
    sensitivity-specificity trade-off at all possible thresholds and evaluates the model without being tied to a single
    threshold. In sports science it is the standard tool for measuring a talent/elite-selection score's discriminative
    power and selecting the best decision threshold.

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss_type_distribution.xlsx
  >> SCENARIO (narration):
    We measure a classification model's discrimination with TSS. For a fair measure
    under class imbalance, TSS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: elite_occurs
    - Predictor(s): ability_index
    - Predictor(s): stress_index

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.26 (acceptable)   Sensitivity = 0.69   Specificity = 0.57   N = 200
    binary outcome: elite_occurs   predictors: ability_index, stress_index

>> COMMENTARY (narration):
    We measured a classification model's (elite occurs/not) discrimination with TSS: TSS = 0.26 (sensitivity 0.69,
    specificity 0.57). TSS = sensitivity + specificity - 1; it excludes chance-expected success and gives a fair
    performance measure even with imbalanced classes. While accuracy can be biased (it anchors to the majority class),
    TSS evaluates the power to capture both positives and negatives together. In sports science it is robust for
    reporting the true discriminative power of talent/elite classification models.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity.xlsx
  >> SCENARIO (narration):
    We evaluate a classifier's predictions against true labels. For detailed
    performance, the confusion matrix is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: actual_label
    - 2nd measure / group: prediction_label

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.93   F1 = 0.918   (true label vs predicted label)

>> COMMENTARY (narration):
    We evaluated a classifier's predictions against true labels via a confusion matrix: accuracy 93%, F1 = 0.92. The
    confusion matrix gathers true/false positives and negatives in one table, from which sensitivity, specificity,
    precision and F1 are derived. Because a single accuracy number can mislead (especially with imbalanced classes),
    this metric set shows where the model errs. In sports science it is fundamental for detailed reporting of
    prediction/classification model performance.

====================================================================================

#63  Random Forest
    file: 63_rf_tree_type.xlsx
  >> SCENARIO (narration):
    We classify discipline from five features with a random forest. For nonlinear
    prediction, random forest is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: discipline
    - Predictor(s): height_cm
    - Predictor(s): weight_kg
    - Predictor(s): VO2_max
    - Predictor(s): sprint_100m
    - Predictor(s): flexibility_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: discipline   predictors: height_cm, weight_kg, VO2_max, sprint_100m, flexibility_score

>> COMMENTARY (narration):
    We classified discipline from five features with a random forest: accuracy 100% -- disciplines are perfectly
    separable with these measures (very high accuracy should be confirmed against overfitting via cross-validation).
    Random forest combines the votes of hundreds of decision trees; it automatically captures nonlinear relationships
    and interactions and also gives a variable-importance ranking. In sports science it is widely used for complex
    classification problems (talent-discipline matching) because it offers both high accuracy and "which variable
    matters" information.

====================================================================================

#64  Support Vector Machines (SVM)
    file: 64_svm_health.xlsx
  >> SCENARIO (narration):
    We classify elite/non-elite from four features with SVM. For well-separable
    classes, SVM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: elite
    - Predictor(s): VO2_max
    - Predictor(s): force_kg
    - Predictor(s): flexibility_cm
    - Predictor(s): balance_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: elite   predictors: VO2_max, force_kg, flexibility_cm, balance_score

>> COMMENTARY (narration):
    We classified elite/non-elite from four features with SVM: accuracy 100% -- the classes are perfectly separable
    with these features (very high accuracy should be confirmed against overfitting via cross-validation). SVM finds
    the decision boundary separating classes with the widest margin; with the kernel trick it can also do nonlinear
    separation. In sports science it is a strong, stable classifier for well-separable class problems, especially on
    small-to-medium data.

====================================================================================

#65  Gradient Boosting
    file: 65_gradient_boosting.xlsx
  >> SCENARIO (narration):
    We classify injury risk from four predictors with gradient boosting. For top
    accuracy, gradient boosting is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: injury_risk
    - Predictor(s): age
    - Predictor(s): late_injury
    - Predictor(s): training_density
    - Predictor(s): BMI

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.929   outcome: injury_risk   predictors: age, late_injury, training_density, BMI

>> COMMENTARY (narration):
    We classified injury risk (injury_risk) from four predictors with gradient boosting: accuracy 92.9%. Gradient
    boosting adds weak trees sequentially -- each new tree corrects the previous model's errors; this is why it wins
    most prediction competitions. While random forest votes in parallel, boosting reduces error step by step. In sports
    science it is preferred where the highest predictive accuracy is sought (injury-risk scoring, outcome prediction);
    it requires tuning against overfitting.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_4kume.xlsx
  >> SCENARIO (narration):
    We cluster athletes by two features. For athlete profiling, k-means is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: VO2_max
    - Variables: force_ratio
    - Number of clusters: 4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.541   variables: VO2_max, force_ratio

>> COMMENTARY (narration):
    We split athletes into 4 clusters by two features (VO2max, force ratio): silhouette = 0.54, a good separation.
    K-means assigns observations to the nearest cluster center and iteratively updates the centers; it groups similar
    athletes into natural clusters. No labels are needed (unsupervised). In sports science it is the core method for
    athlete profiling (endurance/strength types) and forming training groups; silhouette checks the appropriateness of
    the cluster count.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5tur.xlsx
  >> SCENARIO (narration):
    We hierarchically cluster athletes by ten features. For nested group structure,
    hierarchical clustering is appropriate.
  >> VARIABLE SELECTION:
    - Variables: feat_01..10
    - Number of clusters: 5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 5   Silhouette = 0.114   variables: feat_01..10 (10 features)

>> COMMENTARY (narration):
    We split athletes into 5 groups by ten features with hierarchical clustering (silhouette = 0.11, weak separation --
    groups partly overlap). Hierarchical clustering merges observations step by step to build a tree (dendrogram); its
    difference from k-means is not having to fix the cluster count in advance and seeing the nested structure. In sports
    science it is used to explore a hierarchy of athlete groups (main group -> sub-group) and to read the natural
    cluster count from the dendrogram.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan.xlsx
  >> SCENARIO (narration):
    We cluster facility locations with density-based DBSCAN. For spatial clusters
    and outliers, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Variables: lat
    - Variables: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.928   variables: lat, lon (facility location)

>> COMMENTARY (narration):
    We clustered facility/venue locations (lat, lon) with density-based DBSCAN: 4 dense clusters, silhouette = 0.93,
    very good separation. Unlike k-means, DBSCAN does not require the cluster count in advance, can find clusters of any
    shape, and marks sparse points as "noise". In sports science/sports geography it is ideal for detecting spatial
    concentrations (facility clusters, training zones) and isolating outlier locations.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca_6ozellik.xlsx
  >> SCENARIO (narration):
    We reduce six correlated measures to a few components with PCA. For
    dimensionality reduction, PCA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: VO2
    - Variables: force
    - Variables: flexibility
    - Variables: balance
    - Variables: reaction
    - Variables: performance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 explains 93.4% of variance   variables: VO2, force, flexibility, balance, reaction, performance

>> COMMENTARY (narration):
    We reduced six correlated measures to a few components with PCA: the first component alone explains 93.4% of
    variance -- so these six measures largely reflect a single latent dimension (general athletic capacity). PCA
    transforms correlated variables into mutually independent components; it reduces dimensions, eases visualization and
    resolves multicollinearity. In sports science it is fundamental for summarizing multi-measure performance and
    building an "athletic capacity index".

====================================================================================

#70  t-SNE
    file: 70_tsne_5tur.xlsx
  >> SCENARIO (narration):
    We reduce 12-sensor-dimensional data to 2 dimensions with t-SNE for
    visualization. For nonlinear cluster discovery, t-SNE is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sensor_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence = 0.518 (good)   high-dimensional 12 sensors embedded into 2 dimensions

>> COMMENTARY (narration):
    We reduced 12-sensor-dimensional data to 2 dimensions for visualization with t-SNE (KL = 0.52, good quality). t-SNE
    tries to preserve high-dimensional neighborhoods, placing similar observations near and dissimilar ones far; it
    reveals nonlinear cluster structures visually that PCA misses. Interpretation is visual (the axes have no absolute
    meaning). In sports science it is used for exploratory visualization of hidden athlete clusters in high-dimensional
    sensor/measurement data.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds_3grup.xlsx
  >> SCENARIO (narration):
    We place the inter-athlete similarity structure onto a 2D map. For similarity
    maps, MDS is appropriate.
  >> VARIABLE SELECTION:
    - Variables: VO2
    - Variables: flexibility
    - Variables: sprint
    - Variables: balance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.075 (acceptable)   variables: VO2, flexibility, sprint, balance

>> COMMENTARY (narration):
    We placed the distance/similarity structure among athletes onto a 2-dimensional map with low stress (0.08): low
    stress means the map represents the true distances well. MDS positions observations in an interpretable space by
    preserving inter-observation distances; its difference from t-SNE is the aim of preserving global distance
    structure. In sports science it is used to draw athlete/group similarity maps and for positioning between groups.

====================================================================================

#72  UMAP
    file: 72_umap_5tip.xlsx
  >> SCENARIO (narration):
    We reduce 25-dimensional feature data to 2 dimensions with UMAP. For both local
    and global structure, UMAP is appropriate.
  >> VARIABLE SELECTION:
    - Variables: feat_01..25

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    25 features embedded into 2 dimensions   (local-global balance via n_neighbors / min_dist)

>> COMMENTARY (narration):
    We reduced 25-dimensional feature data to 2 dimensions with UMAP. UMAP does nonlinear dimensionality reduction like
    t-SNE but preserves both local and global structure better and is faster. Larger n_neighbors emphasizes broader
    groups, larger min_dist emphasizes the gaps between clusters. In sports science it is a modern choice for
    visualizing high-dimensional performance/measurement data and exploring natural cluster structure.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_21.xlsx
  >> SCENARIO (narration):
    We measure the internal consistency of a 21-item scale. For scale reliability,
    Cronbach's alpha is appropriate.
  >> VARIABLE SELECTION:
    - Variables: item_01..21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach alpha = 0.970 (excellent internal consistency)   21 items

>> COMMENTARY (narration):
    We measured the internal consistency of a 21-item scale with Cronbach's alpha: alpha = 0.97, excellent -- the items
    measure the same construct consistently. Cronbach's alpha shows how much the items of a scale "move together"; above
    0.70 is considered acceptable. (A very high value can also signal item redundancy.) In sports science it is the
    standard index for reporting the reliability of athlete attitude/motivation scales.

====================================================================================

#74  Likert Scale Analysis
    file: 74_likert_3boyut.xlsx
  >> SCENARIO (narration):
    We analyze a 15-item Likert set of three sub-scales
    (motivation/concentration/...). For ordinal scale summary and reliability,
    Likert analysis is appropriate.
  >> VARIABLE SELECTION:
    - Variables: motivation_1..5 / concentration_1..5 / ..._1..5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of items k = 15   Cronbach alpha = 0.757 (acceptable)

>> COMMENTARY (narration):
    We analyzed a 15-item Likert set of three sub-scales (motivation, concentration, ...): reliability is acceptable
    (alpha = 0.76), and item distributions and central tendencies were reported. Likert analysis describes ordinal
    scale responses (1-5) with appropriate summaries (median, distribution, pile-up) and checks scale reliability. In
    sports science it is used for correctly summarizing and interpreting sport-psychology scales (motivation,
    concentration).

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18.xlsx
  >> SCENARIO (narration):
    We discover the latent factors behind eighteen items. For scale-structure
    discovery, EFA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: q01..18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO = 0.905 (excellent)   factor analysis appropriate   18 items (q01..18)

>> COMMENTARY (narration):
    We applied EFA to discover how many latent factors lie behind eighteen items: KMO = 0.91, the data is very suitable
    for factor analysis. EFA reduces observed items to a few unobserved "factors"; it reveals which items measure the
    same dimension. In sports science it is fundamental when developing a new scale (discovering structure) and mapping
    item groups to theoretical dimensions; KMO and Bartlett are prerequisite tests.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3uzman.xlsx
  >> SCENARIO (narration):
    We measure the consistency of scores from three coaches. For inter-rater
    reliability on continuous scores, ICC is appropriate.
  >> VARIABLE SELECTION:
    - Variables: coach_a
    - Variables: coach_b
    - Variables: coach_c

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.903 (Excellent)   95% CI [0.850, 0.940]   3 coaches (coach_a/b/c)

>> COMMENTARY (narration):
    We measured the consistency of scores three coaches gave to the same athletes with ICC: ICC = 0.90, excellent --
    the coaches score almost identically. Unlike kappa, ICC measures inter-rater reliability on CONTINUOUS scores and
    can assess both consistency and absolute agreement. In sports science it is the standard index for determining how
    reliable/interchangeable different coaches'/referees' scores are.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_12_3.xlsx
  >> SCENARIO (narration):
    We test whether a predefined three-factor structure
    (performance/social/motivation) fits the data. For construct validity, CFA is
    appropriate.
  >> VARIABLE SELECTION:
    - Value: performance: performance_1..4
    - Value: social: social_1..4
    - Value: motivation: motivation_1..4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.996   RMSEA = 0.021   3 factors: performance / social / motivation (4 items each)

>> COMMENTARY (narration):
    Unlike EFA, we TESTED whether a pre-defined three-factor structure (performance/social/motivation) fits the data
    with CFA: the fit indices are very good (CFI = 1.00, RMSEA = 0.02). CFA tests a theoretical scale model -- which item
    loads on which factor is fixed in advance, and the question is "does the model fit the data?". In sports science it
    is a mandatory step for confirming the construct validity of a developed scale (do the psychological dimensions
    match theory).

====================================================================================

#78  Survey Mean
    file: 78_survey_means.xlsx
  >> SCENARIO (narration):
    We estimate the mean performance in a stratified/weighted survey. For a complex
    sample, design-based mean is appropriate.
  >> VARIABLE SELECTION:
    - Variables: performance
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    performance: M_hat = 70.93   SE = 0.56   95% CI [69.82, 72.04]   CV 0.79%   weight: weight, stratum: region

>> COMMENTARY (narration):
    In a stratified/weighted survey design we estimated the mean performance while accounting for the design: M = 70.93,
    with a Taylor-linearization SE and 95% CI [69.8, 72.0]. Complex-sample methods account for unequal selection
    probabilities (weights) and stratification; ignoring these biases the standard errors. In sports science it is the
    right way to produce correct point estimates and confidence intervals from national/regional representative athlete
    surveys.

====================================================================================

#79  Survey Frequency
    file: 79_survey_freq.xlsx
  >> SCENARIO (narration):
    We estimate sport-frequency category proportions accounting for the design. For
    a complex sample, design-based frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sport_frequency
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    weighted proportions + Taylor SE + 95% CI for sport_frequency categories
    (e.g. 'low' p_hat = 0.389, 'high' p_hat = 0.048)

>> COMMENTARY (narration):
    We estimated the category proportions of a categorical survey question (sport frequency) under sampling weights and
    stratification; each proportion has a design-based SE and confidence interval. Complex-survey frequency analysis,
    unlike a simple percentage, estimates population proportions without bias by accounting for the sampling design. In
    sports science it is used to correctly report frequency/participation distributions from representative surveys.

====================================================================================

#80  Survey Total
    file: 80_survey_total.xlsx
  >> SCENARIO (narration):
    We estimate the population total (total members) from the sample. For a complex
    sample, design-based total is appropriate.
  >> VARIABLE SELECTION:
    - Variables: member_count
    - Weight: weight
    - Stratum: stratum

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    member_count: T_hat = 267,786   SE = 3,441.17   95% CI [261,012, 274,560]   stratum: stratum

>> COMMENTARY (narration):
    We estimated the population TOTAL (total club members) from the sample: T = 267,786, 95% CI [261,012, 274,560].
    Weights tell how many population units each observation represents; the total is estimated by summing those
    weights, with uncertainty reported via Taylor SE. In sports science it is the right method for producing
    population-scaled totals (total members, total licensed athletes) from a sample.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg.xlsx
  >> SCENARIO (narration):
    We regress performance accounting for the design. For relationships in complex
    samples, design-based regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): age
    - Predictor(s): education_year
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.373   age coef = 0.347 (p < .001)   outcome: performance   predictors: age, education_year   weight: weight

>> COMMENTARY (narration):
    We regressed an outcome (performance) while accounting for the sampling design (weight, stratum): age is a
    significant predictor (b = 0.35, p < .001), and the model explains 37% of variance. Design-based regression
    incorporates weights and the cluster/stratum structure into coefficient and standard-error calculation; ordinary
    regression ignores these and gives biased inference. In sports science it is the right way to model relationships in
    representative athlete survey data.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic.xlsx
  >> SCENARIO (narration):
    We model elite level with design-weighted logistic. For binary outcomes in
    complex samples, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: elite_level
    - Predictor(s): age
    - Predictor(s): motivation
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    age: OR = 0.998   p = 0.853 ns   binary outcome: elite_level   predictors: age, motivation   weight: weight

>> COMMENTARY (narration):
    We modeled a binary survey outcome (elite level or not) with design-weighted logistic regression: the age effect is
    non-significant (OR = 0.998, p = 0.85). This method extends logistic regression to complex sample designs -- weight
    and stratum are reflected in the standard errors. Here the "no significant effect" result is also valuable. In
    sports science it is used to correctly estimate the probability of a binary outcome (elite/non-elite) from
    representative surveys.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_temperature_yield.xlsx
  >> SCENARIO (narration):
    We model performance with age via a flexible curve. For a nonlinear effect, GAM
    is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: performance
    - Predictor(s): age (smooth)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 (explained) = 0.375   outcome: performance   predictor: age (smooth term)

>> COMMENTARY (narration):
    We modeled performance with age, without assuming a straight line in advance, via a flexible curve (smooth): the
    explained variance is moderate (pseudo R^2 = 0.38). GAM extends linear regression -- it models each predictor's
    effect as a smooth function learned from the data, capturing curved relationships without losing interpretability.
    In sports science it is a more explanatory choice than black-box models for relationships where the effect is
    nonlinear (performance that rises with age then declines -- a peak).

====================================================================================

#84  Discriminant Analysis
    file: 84_diskriminant_3sinif.xlsx
  >> SCENARIO (narration):
    We classify athlete class from five continuous measures. To assign to predefined
    groups, discriminant analysis is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: class
    - Predictor(s): sprint
    - Predictor(s): VO2
    - Predictor(s): flexibility
    - Predictor(s): force
    - Predictor(s): reaction

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: class   predictors: sprint, VO2, flexibility, force, reaction

>> COMMENTARY (narration):
    We classified athlete class from five continuous measures with discriminant analysis: accuracy 100% -- the classes
    are fully separable with these measures (overfitting risk should be checked via cross-validation). Discriminant
    analysis finds the linear combinations that best separate groups; it both classifies and shows which variable is
    most influential in separation. In sports science it is used to assign new athletes to predefined groups
    (level/discipline) and to identify discriminating features.

====================================================================================

#85  Conditional Logit
    file: 85_conditional_logit.xlsx
  >> SCENARIO (narration):
    We model athletes' choices among facility alternatives by option features. For
    discrete-choice data, conditional logit is appropriate.
  >> VARIABLE SELECTION:
    - Chooser: athlete_id
    - Choice: chosen
    - Alternative features: fee
    - Alternative features: distance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McFadden pseudo R^2 = 1.00   chooser: athlete_id   choice: chosen   features: fee, distance

>> COMMENTARY (narration):
    We modeled the choices athletes made among facility alternatives by the alternatives' features (fee, distance) with
    a conditional logit. This model is for "discrete choice" data where each individual picks one from a choice set; it
    estimates how an option's features affect its probability of being chosen. Its difference from standard logistic is
    that the choice is conditional on the individual's option set. In sports science it is the core method for
    facility/program preference modeling (fee-distance effect).

====================================================================================

#86  Kaplan-Meier Survival
    file: 86_km_tree.xlsx
  >> SCENARIO (narration):
    We examine athletes' career time to an event and group differences. For censored
    time data, Kaplan-Meier is appropriate.
  >> VARIABLE SELECTION:
    - Time: career_year
    - Event: retired
    - Grouping (categorical): discipline

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 7.79 years   Log-rank chi^2(1) = 0.0003, p = 0.986 ns   time: career_year, event: retired, group: discipline

>> COMMENTARY (narration):
    We examined the career time athletes take to reach an event (retired) with Kaplan-Meier: median survival 7.79
    years; discipline survival curves were compared with the log-rank test (p = 0.99, no difference). KM correctly
    handles censored time data (those whose event has not yet occurred / still active); a plain average ignores these
    observations and is biased. In sports science it is fundamental for career-length, time-to-injury and retention
    analysis.

====================================================================================

#87  Cox Proportional Hazards
    file: 87_cox_hazard.xlsx
  >> SCENARIO (narration):
    We examine retirement risk with continuous and categorical predictors in a Cox
    model. For multiple predictors in censored time, Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: career_year
    - Event: retired
    - Predictor(s): age
    - Predictor(s): past_injury
    - Predictor(s): regular_training

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.556   predictors (HR): age=0.997 (p=0.84), past_injury, regular_training
    time: career_year, event: retired   (predictors not significant alone)

>> COMMENTARY (narration):
    We examined retirement risk with continuous and categorical predictors (age, past injury, regular training) in a
    Cox model: in this sample no predictor was significant alone (age: HR = 0.997, p = 0.84), concordance 0.56
    (weak-moderate discrimination). Cox regression gives the effect of several predictors on "event time" as hazard
    ratios (HR) in censored time data, making no assumption about the baseline hazard's shape. In sports science it is
    the gold standard for identifying the factors that drive career-end/injury risk; here no decisive factor was found.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT)
    file: 88_aft_weibull.xlsx
  >> SCENARIO (narration):
    We model survival time assuming a Weibull distribution. For explicit time
    estimation, AFT is appropriate.
  >> VARIABLE SELECTION:
    - Time: career_year
    - Event: retired
    - Predictor(s): type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution: Weibull   Median survival = 12.71 years   time: career_year, event: retired, predictor: type

>> COMMENTARY (narration):
    We modeled survival time by assuming a parametric (Weibull) distribution: median survival ~12.7 years. AFT
    (accelerated failure time) models, unlike Cox, choose an explicit distribution for the hazard shape and directly
    interpret how predictors "accelerate/decelerate" time. If the data fit the assumed distribution they are more
    powerful than Cox. In sports science they are preferred when explicit career-length estimation and extrapolation
    are needed.

====================================================================================

#89  Competing Risks
    file: 89_competing_risks.xlsx
  >> SCENARIO (narration):
    We model a setting where an athlete can meet several distinct ends. For mutually
    exclusive events, competing risks is appropriate.
  >> VARIABLE SELECTION:
    - Time: career_year
    - Event: retired_cause -> olay tipi / event type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 64   different retirement causes (retired_cause) modeled as distinct event types

>> COMMENTARY (narration):
    We examined a setting where an athlete can meet more than one distinct end (retirement causes: injury, transfer,
    voluntary) in the competing-risks framework: 64 observations censored, the rest split into distinct event types.
    While standard survival treats all events alike, the competing-risks method accounts for the fact that "once one
    occurs the others no longer can" and gives a separate cumulative incidence for each event type. In sports science it
    is necessary to correctly model an athlete's mutually exclusive distinct career ends.

====================================================================================

#90  Time-Dependent Cox
    file: 90_tvcox.xlsx
  >> SCENARIO (narration):
    We build a Cox model where the predictor (stress score) changes over time. For a
    time-varying covariate, time-dependent Cox is appropriate.
  >> VARIABLE SELECTION:
    - Unit (id): athlete_id
    - Start: start
    - Stop: end
    - Event: event
    - Predictor(s): stress_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    stress_score: HR = 0.983   p = 0.289 ns   (time-varying covariate)   id: athlete_id, start/stop intervals

>> COMMENTARY (narration):
    We built a Cox model where the predictor (stress score) CHANGES over time: for each athlete the current stress
    value was used via start-stop intervals; the effect here is non-significant (HR = 0.983, p = 0.29). Time-dependent
    Cox correctly handles non-constant covariates (changing stress, changing conditioning) -- it uses the predictor's
    current value at the event time. In sports science it is the right method for modeling the effect of time-varying
    risk factors (changing stress, load) on career-end/injury.

====================================================================================

#91  Survey Cox Regression
    file: 91_survey_phreg.xlsx
  >> SCENARIO (narration):
    We carry career/survival analysis into a complex survey design (weight+cluster).
    For design-faithful event time, survey_phreg is appropriate.
  >> VARIABLE SELECTION:
    - Time: career_month
    - Event: event
    - Predictor(s): region (faktorize/factorized)
    - Weight: weight
    - Cluster: club_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.516   region_n: p = 0.850 ns   time: career_month, event: event   weight: weight, cluster: club_id

>> COMMENTARY (narration):
    We carried career/survival analysis into a complex survey design (weight + cluster) with a Cox model: the region
    effect is non-significant (p = 0.85), concordance 0.52 (weak discrimination). survey_phreg extends the
    proportional-hazards model to a stratum/cluster/weight structure -- standard errors are corrected for the design.
    In sports science it is the design-faithful way to model event time (career end) in representative panel/survey data.

====================================================================================

#92  Interval-Censored Survival
    file: 92_interval_censored.xlsx
  >> SCENARIO (narration):
    We model data where the event is known only within an interval. For interval
    censoring, this is appropriate.
  >> VARIABLE SELECTION:
    - Lower bound: left_censor
    - Upper bound: survival_censor

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of events = 29   Median survival = 8.0   (lower bound: left_censor, upper bound: survival_censor)

>> COMMENTARY (narration):
    We modeled data where the exact event time is unknown and only known to have occurred within an INTERVAL: median
    survival 8 units, 29 events. Interval-censored methods are for cases where the event lies "somewhere between two
    observations" (between periodic checks); fixing the event to the interval's mid/end point biases results, while this
    method carries the uncertainty correctly. In sports science it is used to correctly model events occurring between
    periodic assessments (injury between two tests).

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty_cox.xlsx
  >> SCENARIO (narration):
    We model the club-specific hidden risk as frailty in clustered survival data.
    For shared hidden risk, frailty Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: career_year
    - Event: event
    - Predictor(s): clinical_group (faktorize/factorized)
    - Cluster: club_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.535   cg_n: p = 0.340 ns   time: career_year, event: event   cluster: club_id (frailty)

>> COMMENTARY (narration):
    In clustered survival data (athletes within the same club) we modeled the club-specific unobserved risk as
    "frailty" (a random effect); the clinical-group effect is non-significant (p = 0.34). Frailty Cox accounts for the
    hidden risk shared by units in the same cluster -- solving the independence assumption that standard Cox violates.
    In sports science it is the right choice for event data clustered within a club/team (shared hidden risk).

====================================================================================

#94  Time Series Analysis
    file: 94_ts.xlsx
  >> SCENARIO (narration):
    We examine a monthly record series (trend, season, stationarity). For
    time-dependent structure, time series analysis is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: monthly_record

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.969 (not stationary)   Trend: increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly record series (monthly_record): the ADF test shows it is not stationary (p = 0.97), with an
    increasing trend and seasonality. Time series analysis reveals the time-dependent structure (trend, season,
    autocorrelation) of observations; ordinary statistics are misleading because of this dependency. Stationarity is a
    prerequisite for models like ARIMA; if non-stationary, differencing is needed. In sports science it is the starting
    step for analyzing performance/participation series.

====================================================================================

#95  STL Decomposition
    file: 95_stl.xlsx
  >> SCENARIO (narration):
    We decompose the mean-performance series into trend, season and residual. For
    seasonal decomposition, STL is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: mean_performance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Season period = 7   series decomposed into trend + season + residual components

>> COMMENTARY (narration):
    We decomposed the mean-performance series into three components with STL: trend, the seasonal pattern (period = 7),
    and residual. STL visually and numerically separates a time series' long-term trend, recurring seasonal pattern and
    unexplained fluctuation; this answers "what is the underlying trend, and how much is seasonal?". In sports science
    it is fundamental for de-seasonalizing performance/participation series (season effect) to see the underlying trend
    and for anomaly detection.

====================================================================================

#96  ARIMA Forecast
    file: 96_arima.xlsx
  >> SCENARIO (narration):
    We model the sport-budget series with ARIMA and produce a forecast. For series
    forecasting, ARIMA is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: sport_budget

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model selected   AIC = 2426.02   forecast + confidence interval

>> COMMENTARY (narration):
    We modeled the sport-budget series (sport_budget) with ARIMA and produced a forecast (AIC = 2426.02 for model
    selection). ARIMA forecasts the future from the series' own past values (AR), trend (I - differencing) and past
    errors (MA); on a stationarized series it gives strong short-to-medium-term forecasts. In sports science/management
    it is the most common classical method for budget/participation/demand forecasting; the confidence interval shows
    the forecast uncertainty.

====================================================================================

#97  Exponential Smoothing (ETS)
    file: 97_ets.xlsx
  >> SCENARIO (narration):
    We model the member-count series with Holt-Winters exponential smoothing. For a
    seasonal trended series, ETS is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: member_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method: Holt-Winters (seasonal additive, period = 12)   AIC = 821.81

>> COMMENTARY (narration):
    We modeled the member-count series (member_count) with Holt-Winters exponential smoothing: it jointly estimates the
    level, trend and a 12-period seasonal component. ETS tracks the series' current level, trend and season by weighting
    recent observations more (exponentially decaying weights); for seasonal and trended series it is a practical
    alternative to ARIMA. In sports science/management it gives fast, reliable results for member/participation
    forecasting with regular seasonal patterns.

====================================================================================

#98  Mann-Kendall Trend
    file: 98_mann_kendall.xlsx
  >> SCENARIO (narration):
    We test a significant trend in the national-performance series without
    distributional assumptions. For robust trend detection, Mann-Kendall is
    appropriate.
  >> VARIABLE SELECTION:
    - Value: national_performance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   significant trend detected   (trend magnitude via Sen's slope)   variable: national_performance

>> COMMENTARY (narration):
    We tested whether the national-performance series has a significant trend -- without assuming a distribution -- with
    Mann-Kendall: p < .001, a significant trend; Sen's slope gives the robust (outlier-resistant) magnitude of the
    trend. Mann-Kendall is a nonparametric trend test; it requires no normality and is resistant to outliers, hence
    common in time-series work. In sports science it is used to robustly detect long-term performance trends
    (rising/falling success trend).

====================================================================================

#99  Anomaly Detection
    file: 99_anomali.xlsx
  >> SCENARIO (narration):
    We detect unusual observations in multivariate data. For composite outlier
    detection, anomaly detection is appropriate.
  >> VARIABLE SELECTION:
    - Variables: VO2_max
    - Variables: force
    - Variables: training
    - Variables: performance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 16   variables: VO2_max, force, training, performance

>> COMMENTARY (narration):
    We automatically detected unusual observations in multivariate data: 16 athletes were flagged as anomalies. Anomaly
    detection finds observations that deviate markedly from normal (erroneous record, exceptional/hidden talent) by
    evaluating several variables together; it catches "composite" outliers that univariate thresholds miss. In sports
    science it is used for data-quality control, erroneous-test detection and the early detection of exceptional athlete
    profiles.

====================================================================================

#100  Variance Components
    file: 100_varcomp_3seviye_h2.xlsx
  >> SCENARIO (narration):
    We decompose measurement variability into nested levels (upper/lower unit). For
    hierarchical variability, variance components is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): upper_unit
    - Factor (categorical): lower_unit (nested)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Random factors: upper_unit + lower_unit (nested within upper)   % contribution of each component + residual

>> COMMENTARY (narration):
    We decomposed a measurement's variability into nested levels -- upper unit and lower unit within upper: the percent
    contribution of each level and the residual were reported. Variance components analysis answers "how much of the
    variability is between upper units, how much between lower units, how much within unit?". In sports science it is
    used in hierarchical structures (club>team>athlete) to see where uncertainty concentrates and in sampling design.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old.xlsx
  >> SCENARIO (narration):
    We examine two groups' measurement difference with a Bayesian t-test, expressing
    evidence as a Bayes factor. For an intuitive evidence ratio, this is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Grouping (categorical): group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 3.36e+46   Cohen's d = 3.939   (overwhelming evidence for H1)   outcome: value, group: group

>> COMMENTARY (narration):
    We examined the measurement difference between two groups (new/old method) with a Bayesian t-test: the Bayes factor
    is overwhelming (BF10 ~ 3.4e46), the effect very large (d = 3.94) -- the data support the "difference" hypothesis
    over "no difference" by astronomical odds. Unlike a p-value, the Bayes factor gives the RELATIVE evidence strength
    of two hypotheses and can distinguish "no evidence" from "no difference". In sports science it is preferred when one
    wants to express the evidential strength of a decision as an intuitive ratio.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BF10.xlsx
  >> SCENARIO (narration):
    We evaluate the relationship between two variables in a Bayesian framework. For
    evidential strength, Bayesian correlation is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: X_variable
    - 2nd measure / group: Y_variable

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.780 (very strong)   BF10 = 4.50e+14   (overwhelming evidence for a relationship)

>> COMMENTARY (narration):
    We evaluated the relationship between two variables in a Bayesian framework: r = 0.78 and BF10 ~ 4.5e14, i.e. very
    strong evidence for a relationship. Bayesian correlation, instead of a classical p-value, presents the evidential
    strength of the relationship as a Bayes factor and the coefficient's posterior distribution. In sports science it is
    valuable for reporting not just whether the relationship between two indicators is "significant" but how strongly it
    is "evidenced".

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_2yonlu.xlsx
  >> SCENARIO (narration):
    We examine a measurement's difference across two factors with Bayesian ANOVA.
    For the evidential strength of factor effects, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): factor1
    - 2nd Factor: factor2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor1: BF10 = 76099 (very strong evidence)   outcome: value, factors: factor1, factor2

>> COMMENTARY (narration):
    We examined a measurement's difference across two factors with Bayesian ANOVA: for factor1, BF10 ~ 76,000, very
    strong evidence. Bayesian ANOVA compares the effects of factors and their interactions via Bayes factors; it ranks
    probabilistically which model (which effects) best explains the data. Unlike classical ANOVA's "reject/don't reject"
    decision, it quantifies the relative support among models. In sports science it is used to compare the evidential
    strength of factor effects.

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_LMM.xlsx
  >> SCENARIO (narration):
    We analyze group-nested data with a Bayesian hierarchical model. For stable
    estimates in small groups, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_response
    - Cluster: group_id
    - Predictor(s): X_covariate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2_u (group) = 2.90 (8.6%)   sigma^2_eps (residual) = 30.79 (91.4%)   outcome: Y_response, group: group_id

>> COMMENTARY (narration):
    We analyzed group-nested data with a Bayesian hierarchical (multilevel) model: 8.6% of variability is
    between-group, 91.4% within-group (ICC ~ 0.09). The Bayesian hierarchical model is the Bayesian version of LMM -- it
    estimates group effects with posterior distributions and balances small groups via "partial pooling". In sports
    science it is powerful for producing stable estimates even in small groups within multilevel (club/team/athlete)
    data.

====================================================================================

#105  Spatial SAR
    file: 105_spatial_sar_spatial.xlsx
  >> SCENARIO (narration):
    When modeling a measurement we handle spatial spillover with SAR. For
    neighborhood effects, SAR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho = 0.336   z = 4.94   p < .001 ***   Pseudo R^2 = 0.87   (N = 100, k-NN W, k = 5)
    outcome: Y_value, predictors: X1, X2

>> COMMENTARY (narration):
    When modeling a measurement (Y_value), we handled spatial spillover (the effect of neighboring units) with SAR: the
    spatial lag parameter is significant and positive (rho = 0.34, p < .001) -- a unit's value is related to its
    neighbors' value, a "cluster/spillover" pattern. SAR incorporates spatial dependency into the model; if ignored,
    standard errors are biased. In sports science/sports geography it is the right method for modeling the geographic
    spread of club/region performance (neighborhood effect).

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_residual.xlsx
  >> SCENARIO (narration):
    We model spatial dependency in the error term. For unmeasured geographic
    factors, SEM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda = 0.590   z = 6.38   p < .001 ***   Pseudo R^2 = 0.84   (N = 100, k-NN W, k = 5)

>> COMMENTARY (narration):
    This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
    0.59, p < .001) -- the effect of geographic variables omitted from the model makes neighboring errors correlated.
    Unlike SAR, SEM attributes the spread to the error rather than the outcome. In sports science/sports geography, when
    the source of spatial autocorrelation is unmeasured geographic factors (infrastructure, access), the correct
    specification is SEM; it is chosen by comparison with SAR.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local.xlsx
  >> SCENARIO (narration):
    Assuming the relationship is not constant in space, we estimate separate
    coefficients per location. For spatial heterogeneity, GWR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.887   local coefficients vary by location   outcome: Y_value, predictors: X1, X2

>> COMMENTARY (narration):
    Assuming the relationship is NOT constant across space, we estimated SEPARATE coefficients for each location (R^2 =
    0.89). Unlike "global" regression, GWR fits a separate model at each point with local overlap/weights; this answers
    "how does this variable's effect vary by region?" and maps spatial heterogeneity. In sports science/sports geography
    it is powerful where the performance-resource relationship differs by region.

====================================================================================

#108  Nested Mixed Model
    file: 108_nested_lmm_R_P_F.xlsx
  >> SCENARIO (narration):
    We model a value in a nested design (upper>lower>block). For nested hierarchies,
    nested LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): lower_group_no
    - Factor (categorical): upper_group
    - Factor (categorical): block

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper_group (P) effect: F = 27.44, p < .001 ***   lower_group_no (R) effect: F = 6.11, p < .001 ***
    nested variance components (block within P)

>> COMMENTARY (narration):
    We modeled a value in a nested design (upper group > lower group > block): both the upper level (F = 27.44, p <
    .001) and the lower/replication level (F = 6.11, p < .001) make significant contributions. Nested LMM correctly
    handles hierarchies where sub-units are nested within super-units (each block belongs to only one group); by
    partitioning variance into levels it shows each layer's share. In sports science it is used to correctly separate
    effects in club>team>athlete hierarchies.

====================================================================================

#109  Crossed Mixed Model
    file: 109_crossed_lmm_A_B.xlsx
  >> SCENARIO (narration):
    We model a design where two random factors are crossed. For two independent
    classification axes, crossed LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): factor_a
    - 2nd Factor: factor_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    A x B interaction: F = 14.67, p < .001 ***   main effects ns (A: p=0.767, B: p=0.480)   outcome: value

>> COMMENTARY (narration):
    We modeled a design where two random factors are crossed rather than nested (each level of A pairs with each level
    of B): while the main effects are non-significant (A: p=0.77, B: p=0.48), the A x B interaction is very strong
    (F = 14.67, p < .001) -- so the effect depends on the COMBINATION of factors. Crossed LMM, unlike nested, handles
    two independent grouping axes (e.g. referee x athlete, each referee on each athlete) at once. In sports science it
    is the right choice for jointly analyzing the effects and interaction of two independent classification axes.

====================================================================================

#110  Kernel Density (KDE) Map
    file: 110_KDE_traffic_density.xlsx
  >> SCENARIO (narration):
    We produce a continuous density surface from unit locations. For a density map,
    KDE is appropriate.
  >> VARIABLE SELECTION:
    - Value: performance_score
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 95 units   performance-weighted spatial density surface (continuous heat map)

>> COMMENTARY (narration):
    Using unit locations (n = 95) we produced a continuous density surface (a KDE heat map): the point distribution was
    turned into a smooth "density" surface. KDE answers "where is it dense?" from scattered point data as a continuous
    map; it shows the trend rather than individual points. In sports science/sports geography it is the core spatial
    tool for visually mapping facility/athlete/performance concentrations and identifying empty/dense zones.

====================================================================================

#111  Hexbin Density Map
    file: 111_Hexbin_measurement_noktalari.xlsx
  >> SCENARIO (narration):
    We aggregate unit distribution into hexagonal cells to show density. For the
    over-plotting problem, hexbin is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 150 units   spatial density via aggregation into hexagonal cells

>> COMMENTARY (narration):
    We colored unit distribution (n = 150) by counting within hexagonal cells. Hexbin aggregates many overlapping
    points (the over-plotting problem) into a regular hexagonal grid to show density clearly; compared with squares it
    carries less directional bias. In sports science/sports geography it is used to turn dense point clouds (facility/
    athlete locations) into a readable density map and to compare spatial concentrations.

====================================================================================

#112  Moran's I
    file: 112_Morans_I_structure_quality_autocorrelation.xlsx
  >> SCENARIO (narration):
    We test whether the performance score is distributed randomly or in clusters
    across space. For spatial autocorrelation, Moran's I is appropriate.
  >> VARIABLE SELECTION:
    - Value: performance_score
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.563   z = 10.13   p < .001 ***   (k-NN neighbors, k = 6)   variable: performance_score

>> COMMENTARY (narration):
    We tested whether the performance score is distributed randomly or in clusters across space with Moran's I: I =
    0.56, z = 10.13, p < .001 -- strong positive spatial autocorrelation, i.e. similar performance values cluster
    geographically. Moran's I quantifies "Tobler's first law" (near things are similar). In sports science/sports
    geography it is the first test for detecting the geographic clustering of performance/participation variables and
    for deciding whether a spatial model is needed.

====================================================================================

#113  Getis-Ord Gi*
    file: 113_Getis_Ord_traffic_hotspot.xlsx
  >> SCENARIO (narration):
    We map locally where the performance score clusters high/low. For hot/cold
    spots, Getis-Ord is appropriate.
  >> VARIABLE SELECTION:
    - Value: performance_score
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Hot spots = 31   Cold spots = 30   n = 68   (k-NN k = 8, binary weights)   variable: performance_score

>> COMMENTARY (narration):
    We mapped locally WHERE the performance score clusters high/low with Getis-Ord Gi*: 31 significant hot spots
    (high-performance clusters) and 30 cold spots (low-performance clusters). While Moran's I states the overall
    clustering, Gi* shows its location -- for each point it tests "is the surrounding area high or low?". In sports
    science/sports policy it is used to pinpoint high/low-performance zones (hot/cold spots) for targeted
    investment/infrastructure.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_structure_kumeleri.xlsx
  >> SCENARIO (narration):
    We cluster unit locations with density-based DBSCAN. For spatial clusters and
    outlier locations, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Clusters = 4   Noise (outliers) = 9   n = 115   (location: lat, lon)

>> COMMENTARY (narration):
    We clustered unit locations (n = 115) with density-based DBSCAN: 4 natural geographic clusters and 9 "noise"
    (scattered, cluster-less) points. DBSCAN does not require the cluster count in advance, finds clusters of any shape,
    and separates sparse points as outliers -- ideal for spatial clustering. In sports science/sports geography it is
    used to detect region-level natural concentrations (facility/athlete clusters) and to isolate isolated locations;
    it completes our descriptive spatial analysis series.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_VO2max_mL_kg_min.xlsx
  >> SCENARIO (narration):
    We follow 40 athletes measured at three time levels (pre, mid, post); each belongs to one of two training groups (HIIT / continuous). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on VO2max.
  >> VARIABLE SELECTION:
    - Dependent variable: VO2max_mL_kg_min
    - Subject ID: athlete_id
    - Between-subjects factor: training
    - Within-subjects factor: time

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (training): F(1,38) = 30.61  p < .001  np2 = 0.446
    Within (time):  F(2,76) = 80.46  p < .001  np2 = 0.679
    Interaction:         F(2,76) = 11.82  p < .001  np2 = 0.237
    Mauchly W = 0.941  p = 0.317   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a training x time mixed design we analyzed VO2max for 40 athletes (120 observations). The interaction is significant (F(2,76) = 11.82, p < .001, np2 = 0.237) *** -- the two groups' change across time differs in magnitude. The between-subjects main effect (HIIT vs continuous) is F = 30.61, p < .001; the within-subjects main effect (pre/mid/post) is F = 80.46, p < .001. Mauchly's test p = 0.317, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each time level. In sport sciences, the mixed design is the standard analysis for comparing training protocols on aerobic capacity over a program.

====================================================================================
