====================================================================================
MerQur - Sosyal_Beseri - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Sosyal_Beseri/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_demografi.xlsx
  >> SCENARIO (narration):
    We compiled the demographic inventory of 280 participants. Age, education years,
    income, satisfaction and weekly study were measured for each. Before any
    inferential test we want the overall picture of the sample; so we begin with
    descriptive statistics.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: education_year
    - Variables: monthly_income_TL
    - Variables: satisfaction
    - Variables: weekly_study
    - Grouping (categorical): type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 participants
    age: mean = 43.9   |   education years: mean = 13.1   |   monthly income: mean = 12,228 TL
    satisfaction: mean = 3.47 (of 5)   |   weekly study: mean = 40.3 hours

>> COMMENTARY (narration):
    First we draw the overall picture of our sample: 280 participants, average age ~44, 13 years of education, mean
    monthly income ~12,200 TL, satisfaction of 3.47 out of 5, and ~40 weekly study hours. This descriptive table lays
    the groundwork for every analysis that follows -- group comparisons, achievement-motivation relationships, spatial
    socioeconomic pattern. In the social sciences and humanities, before any inferential test, summarizing the
    sample's basic features (age, education, income, satisfaction) is essential both to audit data quality and to set
    research priorities.

====================================================================================

#2  Normality Tests
    file: 02_normality_income_motivation.xlsx
  >> SCENARIO (narration):
    We examine whether income, motivation and anxiety are normally distributed.
    Because subsequent t-tests, ANOVA and correlation depend on this assumption, we
    test each variable separately.
  >> VARIABLE SELECTION:
    - Variables: income
    - Variables: motivation
    - Variables: anxiety

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    3 continuous variables tested:
    income                  : Shapiro-Wilk = 0.899  p < .001    KS p = 0.001   Not Normal
    motivation              : Shapiro-Wilk = 0.994  p = 0.657    KS p = 0.636   Normal
    anxiety                 : Shapiro-Wilk = 0.904  p < .001     KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three continuous variables -- income, motivation and anxiety -- are normally distributed. The
    result splits instructively: motivation is normal (p = 0.66), but income and anxiety deviate significantly from
    normality (p < .001). This is typical in social data -- income is right-skewed, and anxiety is an ordinal/bounded
    scale (floor/ceiling effect). Practical upshot: we can safely use parametric tests (t-test, ANOVA, Pearson) on
    motivation; for income and anxiety, nonparametric methods (Mann-Whitney/Kruskal-Wallis) or a transform are more
    appropriate. The normality check is a critical preliminary step that decides, per variable, which test family fits.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_GPA.xlsx
  >> SCENARIO (narration):
    We investigate whether students' mean GPA (4-point scale) differs from a 2.5
    reference threshold. With one group and a fixed reference, the one-sample t-test
    is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: GPA
    - Test value (mu): 2.5

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = 6.242   p < .001 ***   Cohen d = 0.624 (Medium)
    Mean GPA = 2.83   (test mu = 2.5, reference threshold / 4-point scale)   H0 REJECTED

>> COMMENTARY (narration):
    We compared students' mean GPA (4-point scale) against a reference threshold of 2.5. The result is significant and
    medium-sized: mean 2.83, above the threshold -- t(99) = 6.24, p < .001, d = 0.62. Students perform statistically
    above the reference. The one-sample t-test is the right way to compare a measure against a known standard/threshold
    (passing grade, target GPA); in education research it is widely used to evaluate achievement against a benchmark.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent-Samples t-Test
    file: 04_independent_t_sex_score.xlsx
  >> SCENARIO (narration):
    We compare the mean achievement score of two independent groups (sex). With two
    separate groups and a continuous measure, the independent-samples t-test is
    appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): sex
    - Test variable: score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = 3.606   p < .001 ***   Cohen d = 0.673 (Medium)   H0 REJECTED

>> COMMENTARY (narration):
    We compared the mean achievement score of two independent groups (by sex). The difference is significant and
    medium-sized: t(113) = 3.61, p < .001, d = 0.67. It shows the between-group difference is too pronounced to be
    chance and is also practically noteworthy. The independent-samples t-test is the standard way to compare the means
    of two separate groups (male/female, two classes, two schools) on a continuous measure; it is a fundamental tool
    for detecting group differences in education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired-Samples t-Test
    file: 05_paired_t_on_last.xlsx
  >> SCENARIO (narration):
    We compare two scores measured before and after an intervention in the same
    students. Since the measures are paired, the paired t-test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: on_test
    - 2nd measure / group: last_test

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(49) = -15.99   p < .001 ***   Cohen d_z = -2.262 (Large)
    Mean diff (on_test - last_test) = -12.23   H0 REJECTED

>> COMMENTARY (narration):
    We paired and compared two scores measured before and after an intervention in the same students (pre-test vs
    post-test). The result is very strong: mean difference -12.2 points, t(49) = -15.99, p < .001, d_z = -2.26, a huge
    effect. Post-intervention scores rose significantly and substantially. The paired t-test compares two timed
    measures on the same unit (before/after training); by isolating individual change it is more powerful than the
    independent test and is the right way to measure an intervention's effect in education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_instruction.xlsx
  >> SCENARIO (narration):
    We compare the effect of four instruction methods on mean achievement. With more
    than two groups, one-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Factor (categorical): instruction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,136) = 15.38   p < .001 ***   eta^2 = 0.253   H0 REJECTED

>> COMMENTARY (narration):
    We tested the effect of four different instruction methods on mean achievement. The result is significant and
    strong: F(3,136) = 15.38, p < .001, eta^2 = 0.25 -- about a quarter of score variance comes from method
    differences. At least one method differs significantly. One-way ANOVA compares the means of more than two groups at
    once (avoiding the error inflation of many t-tests); in education it is the core method for comparing different
    method/group/class performance. Which pairs differ is then determined by post-hoc tests.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova.xlsx
  >> SCENARIO (narration):
    We examine the main effects and interaction of method and sex on achievement
    simultaneously. With two categorical factors, two-way ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Factor (categorical): method
    - 2nd Factor: sex

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    method (main effect): F(2) = 43.85   p < .001 ***   eta^2p = 0.379
    sex (main effect)   : F(1) = 0.42    p = 0.519 ns    eta^2p = 0.003

>> COMMENTARY (narration):
    We examined two factors at once: how do instruction method and sex affect the achievement score? The method main
    effect is very strong (F(2) = 43.85, p < .001, eta^2p = 0.38), while the sex main effect is non-significant
    (p = 0.52). The decisive driver of the score is method; sex alone contributes nothing notable. The power of two-way
    ANOVA is that it tests both factors and their interaction in a single model -- isolating each factor's pure effect
    with the other controlled. It is ideal for answering "which variable really makes a difference?" in education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated-Measures ANOVA
    file: 08_repeated_anova_term_GPA.xlsx
  >> SCENARIO (narration):
    We compare achievement over four consecutive terms in the same students. With
    repeated measures on the same unit, repeated-measures ANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: measurement_1
    - Repeated measures: measurement_2
    - Repeated measures: measurement_3
    - Repeated measures: measurement_4

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3,177) = 108.40   p < .001 ***   eta^2p = 0.648   n = 60   H0 REJECTED

>> COMMENTARY (narration):
    We compared achievement measured across four consecutive terms in the same 60 students. The result is very strong:
    F(3,177) = 108.40, p < .001, eta^2p = 0.65 -- the between-term difference is huge and most of the effect is
    time-related. Student achievement changes significantly across terms. Repeated-measures ANOVA compares three or
    more timed measurements on the same unit; by holding individual differences constant it yields high statistical
    power and is the right choice for multi-term progress tracking (term GPA, follow-up measures) in education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_method_3ders.xlsx
  >> SCENARIO (narration):
    We test instruction method's effect on math, science and language achievement at
    once. With several correlated dependent variables, MANOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: math
    - Dependent variable: science
    - Dependent variable: language
    - Factor (categorical): method

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.433   F(6,230) = 19.90   p < .001 ***   n = 120
    Dependent: math, science, language   |   Factor: method   H0 REJECTED

>> COMMENTARY (narration):
    We tested instruction method's effect on three dependent variables (math, science, language achievement)
    simultaneously. With Wilks' Lambda = 0.43, F(6,230) = 19.90, p < .001, the effect is very strong. MANOVA examines
    several correlated outcomes in one test, both preventing the error inflation of many separate ANOVAs and capturing
    the joint information the variables carry together. In education, when method affects not a single course but the
    course-achievement "bundle" together, MANOVA reveals this multivariate difference as a single decision.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova.xlsx
  >> SCENARIO (narration):
    We compare post-test scores across groups while controlling the pre-test score
    as a covariate. With a confounding continuous variable, ANCOVA is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: last_test
    - Factor (categorical): group
    - Covariate: on_test

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group effect significant   eta^2p = 0.474 (large)   covariate: on_test   H0 REJECTED

>> COMMENTARY (narration):
    We compared post-test scores (last_test) across groups while controlling the pre-test score (on_test) as a
    covariate. This rules out the objection that "the groups differed at baseline" and measures the pure group effect:
    the effect is large (eta^2p = 0.47). ANCOVA makes the group comparison fair by statistically holding a confounding
    continuous variable constant -- it is the answer to "once we equalize the starting level, does the intervention
    still make a difference?" and is the standard tool in pre-test/post-test designs in education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci_biomass.xlsx
  >> SCENARIO (narration):
    For mean daily absenteeism we build a confidence interval via resampling, with
    no distributional assumption. For skewed data, bootstrap is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: absenteeism_day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean (absenteeism/day) = 4.53
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    For the mean daily absenteeism we produced a 95% confidence interval via resampling -- with no distributional
    assumption: observed mean 4.53 days. Bootstrap builds the sampling distribution of the statistic empirically by
    resampling the data thousands of times from itself; it is a reliable way to give a confidence interval when
    normality does not hold or no formula is known. In education it provides more robust estimates than classic
    t-intervals for skewed quantities (absenteeism, tardiness, disciplinary events).

====================================================================================

#12  Permutation Test
    file: 12_permutation_method_yield.xlsx
  >> SCENARIO (narration):
    We test the motivation difference between two groups via permutation, with no
    distributional assumption. For small samples/odd distributions, permutation is
    appropriate.
  >> VARIABLE SELECTION:
    - Test variable: motivation
    - Grouping (categorical): group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = -2.247   p = 0.029 *   significant mean difference between the two groups   H0 REJECTED

>> COMMENTARY (narration):
    We tested the motivation difference between two groups without any distributional assumption, via permutation: the
    null distribution was built by randomly swapping group labels, and the observed difference turned out rare in that
    distribution (p = 0.029). Because the permutation test is exact and distribution-free, it is a safe alternative to
    the parametric t-test for small samples or odd distributions; in education it is a robust choice for group
    comparisons where assumptions are in doubt.

====================================================================================

#13  Multiple Comparison
    file: 13_multiple_comparison_method.xlsx
  >> SCENARIO (narration):
    We take the p-values of eight material comparisons together and apply
    multiple-testing correction. With many tests, p-adjustment is appropriate.
  >> VARIABLE SELECTION:
    - Variables: material
    - Variables: score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Raw: 5/8 significant   |   After Bonferroni / Holm / BH (FDR): 3/8 significant

>> COMMENTARY (narration):
    We took the p-values of eight separate material/group comparisons together and applied multiple-testing correction.
    Raw, 5 comparisons were significant; after Bonferroni/Holm/BH that dropped to 3. When many tests are run, the rate
    of false positives that look "significant" by chance alone inflates; correction methods tighten the threshold to
    control this error. In education, when many methods/materials are compared at once, some "significant" conclusions
    reached without correction can be misleading -- this step protects inference reliability.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney_agriculture_quality.xlsx
  >> SCENARIO (narration):
    We compare two independent groups on an ordinal measure. Since normality fails,
    Mann-Whitney is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): location
    - Test variable: survival_satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    U = 1121.50   p = 0.049 *   r = -0.25   H0 REJECTED

>> COMMENTARY (narration):
    We compared the distribution of two independent groups (location) on an ordinal/skewed measure
    (survival_satisfaction) based on ranks rather than means: U = 1121.5, p = 0.049, small-medium effect (r = -0.25).
    The difference is borderline significant. Mann-Whitney is the nonparametric counterpart of the t-test; when
    normality fails or the scale is ordinal, it is the right way to compare two groups. In education/social research it
    is a robust choice for ordinal measures like satisfaction/grade.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank
    file: 15_wilcoxon_tree_health.xlsx
  >> SCENARIO (narration):
    We compare attitude scores measured before/after in the same units, without
    assuming normality. For paired non-normal data, Wilcoxon is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: attitude_before
    - 2nd measure / group: attitude_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    W = 0.00   p < .001 ***   r = 0.89   n (non-zero) = 30   H0 REJECTED

>> COMMENTARY (narration):
    We compared attitude scores measured before and after on the same units (attitude_before vs post) -- without
    assuming normality, on a rank basis: W = 0, p < .001, very large effect (r = 0.89). Attitude changed consistently
    and strongly after the intervention. Wilcoxon is the nonparametric counterpart of the paired t-test; it is the
    right choice for ordinal or non-normal before/after measures. In education it gives reliable results for pre/post
    attitude and satisfaction measures.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_soil_yield.xlsx
  >> SCENARIO (narration):
    We compare three groups on a numeric/ordinal measure. Since normality/variance
    homogeneity fails, Kruskal-Wallis is appropriate.
  >> VARIABLE SELECTION:
    - Grouping (categorical): sound
    - Test variable: book_count

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 15.96   p < .001 ***   eta^2_H = 0.33   H0 REJECTED

>> COMMENTARY (narration):
    We compared the distribution of a numeric/ordinal measure (book_count) across more than two groups (sound/noise
    level) by ranks rather than means: H(2) = 15.96, p < .001, eta^2_H = 0.33, a strong effect. At least one group
    differs significantly. Kruskal-Wallis is the nonparametric counterpart of one-way ANOVA; when normality or variance
    homogeneity fails, it is the right way to do multi-group comparison. In education/social data it is robust for
    detecting group differences on skewed or ordinal measures.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman.xlsx
  >> SCENARIO (narration):
    We compare scores measured under four conditions in the same units. As a
    nonparametric repeated measure, Friedman is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: condition_A
    - Repeated measures: condition_B
    - Repeated measures: condition_C
    - Repeated measures: condition_D

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 39.08   p < .001 ***   Kendall W = 0.434   n = 30   H0 REJECTED

>> COMMENTARY (narration):
    We compared scores measured under four conditions (condition_A..D) in the same 30 units -- as a nonparametric
    repeated measure: chi^2(3) = 39.08, p < .001, Kendall W = 0.43 (moderate concordance). The between-condition
    difference is significant. Friedman is the nonparametric counterpart of repeated-measures ANOVA; it is the right
    choice for ordinal or non-normal paired multi-condition measures. In education it is used to compare the same
    student's rankings across different methods/conditions.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_seedling.xlsx
  >> SCENARIO (narration):
    We test the proportion of students who passed against an expected 50%. With a
    binary outcome and a theoretical proportion, the binomial test is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: passed
    - Expected proportion: 0.50
    - Success value: 1

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed proportion = 0.704   (expected p = 0.50)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested the proportion of students who passed against an expected 50%: observed proportion 70.4%, well above
    expectation (p < .001). The pass rate is too high to be chance. The binomial test is the exact method for comparing
    the observed proportion of a binary (pass/fail) outcome with a theoretical proportion; in education it directly
    tests whether success/pass/conversion rates meet a target or a 50:50 expectation.

====================================================================================

#19  Sign Test
    file: 19_sign_test_nitrogen.xlsx
  >> SCENARIO (narration):
    We look at the direction of before/after scores in the same students. When only
    directional information is reliable, the sign test is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: score_before
    - 2nd measure / group: score_post

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive diffs (before > after) = 0   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We looked at the direction of change in scores measured before/after on the same students (score_before vs post):
    all changes go in one direction, none in reverse (positive diffs = 0), p < .001. The sign test uses only the
    direction of the difference (increase/decrease), not its magnitude, so it is the before/after test that requires
    the fewest assumptions. In education it is a robust choice when the measure is skewed or only directional
    information is reliable (did the score go up or down).

====================================================================================

#20  Runs Test
    file: 20_runs_test.xlsx
  >> SCENARIO (narration):
    We test whether the absence sequence (around the median) is random. For sequence
    randomness, the runs test is appropriate.
  >> VARIABLE SELECTION:
    - Column: absent

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 28   Z = -0.243   p = 0.808 ns   data may be considered random

>> COMMENTARY (narration):
    We tested whether the binary sequence of absences (around the median) is random: 28 runs, Z = -0.24, p = 0.81 --
    no pattern, the sequence is random. The runs test checks whether values in a sequence form a systematic pattern
    (clusters, cycles, trend). In education/social series it is used to determine whether systematic patterns or
    randomness dominate, and in checking process control and the independence assumption.

====================================================================================

#21  Chi-Square Independence
    file: 21_chisquare_type_health.xlsx
  >> SCENARIO (narration):
    We test whether noise level and university preference are related. With two
    categorical variables, chi-square is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sound
    - Variables: university

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(2) = 141.06   p < .001 ***   Cramer's V = 0.511 (Strong)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether two categorical variables -- noise level (sound) and university preference (university) -- are
    related: chi^2(2) = 141.06, p < .001, Cramer's V = 0.51, a strong dependency. The categories are not independent;
    they vary together. The chi-square test of independence detects the relationship between two qualitative variables
    from a cross-tab; in education/social research it is the core method for revealing categorical relationships such as
    group-preference, region-attitude. Cramer's V measures the practical strength of the relationship.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_quality.xlsx
  >> SCENARIO (narration):
    We test whether a four-category learning style's observed distribution fits an
    equal expected distribution. For one categorical variable, goodness-of-fit is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: style

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(3) = 22.88   p < .001 ***   N = 200, k = 4 (expected: equal distribution)   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the observed distribution of a four-category variable (style/learning style) fits an equal (1/k)
    expected distribution: chi^2(3) = 22.88, p < .001 -- the categories are not equally distributed, some are clearly
    more frequent. The goodness-of-fit test compares the observed frequencies of a single categorical variable with a
    theoretical expectation (equal proportions, a known ratio); in education it tests whether preference/style
    distributions match an expected profile.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_method.xlsx
  >> SCENARIO (narration):
    In a 2x2 table we test the method-pass relationship exactly, due to small cell
    frequencies. For few observations, Fisher is appropriate.
  >> VARIABLE SELECTION:
    - Variables: method
    - Variables: pass

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher exact p = 0.018 *   Phi = 0.486 (Strong)   H0 REJECTED

>> COMMENTARY (narration):
    In a 2x2 cross-tab (method x pass) we tested the relationship exactly -- using Fisher instead of chi-square because
    of small cell frequencies: p = 0.018, with a strong relationship (Phi = 0.49). Fisher's exact test is the right
    choice when expected frequencies are low and the chi-square approximation is unreliable; it computes the
    probability exactly rather than approximately. In education it gives reliable results for small-sample pilot/trial
    comparisons (does the new method work).

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_test.xlsx
  >> SCENARIO (narration):
    We compare a binary outcome at two times in the same students (pass/fail). For
    paired binary change, McNemar is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: on_test
    - 2nd measure / group: last_test

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2 = 0.00   p = 1.000 ns   discordant b = 13, c = 14   OR = 0.929   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested the direction of change in a binary outcome measured at two times in the same students (on_test vs
    last_test: pass/fail): the changes are nearly balanced (b = 13, c = 14), p = 1.00 -- no significant directional
    change. McNemar tests the before/after binary status change in the same unit and looks only at discordant pairs.
    In education it is the right tool for detecting whether status changes before/after an intervention (pass/fail) are
    systematic.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_expert.xlsx
  >> SCENARIO (narration):
    We measure the agreement of two teachers classifying the same students. For
    categorical agreement, kappa is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: teacher_a
    - 2nd measure / group: teacher_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Kappa = 0.585 (moderate agreement)

>> COMMENTARY (narration):
    We measured the agreement of two teachers (teacher_a vs teacher_b) in assigning the same students to the same
    categories -- excluding the agreement expected by chance: kappa = 0.585, moderate. Unlike raw percent agreement,
    kappa reports categorical agreement after removing the chance-agreement share, so it is more honest. In education
    it is the standard index for measuring how consistent two teachers'/raters' classifications (achievement level,
    performance class) are.

====================================================================================

#26  Cochran-Mantel-Haenszel
    file: 26_cmh_region_method.xlsx
  >> SCENARIO (narration):
    We test the group-success relationship controlling for school strata. For a
    stratum-controlled relationship, CMH is appropriate.
  >> VARIABLE SELECTION:
    - Variables: group
    - Variables: successful
    - Stratum: school

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi^2(1) = 19.67   p < .001 ***   Common OR (MH) = 2.77 [1.76, 4.37]
    Breslow-Day p = 0.982 (OR homogeneous)   H0 REJECTED

>> COMMENTARY (narration):
    We tested the group-success relationship while controlling for school strata: the common Odds Ratio across strata =
    2.77 [1.76, 4.37], p < .001; Breslow-Day p = 0.98 means this relationship is consistent across all strata. CMH
    measures the pure strength of the association by holding a confounding stratum variable constant (preventing
    Simpson's paradox). In education it is ideal for robustly estimating a group-outcome relationship while controlling
    school/region differences.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_type_region_disease.xlsx
  >> SCENARIO (narration):
    We model the joint relationship structure of three categorical variables. For
    more than two categorical dimensions, log-linear is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sex
    - Variables: sound
    - Variables: university_reader

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 46.72   Pearson chi^2 = 0.0 (saturated model)   joint relationship of 3 categorical variables modeled

>> COMMENTARY (narration):
    We examined the joint relationship structure of three categorical variables (sex, noise, university reader) with a
    log-linear model. The model explains cell frequencies via main effects and interactions; it reveals which
    pairs/triples of variables vary together. It is the multivariable generalization of the two-way cross-tab. In
    education/social research it is used to analyze the joint dependency pattern of more than three categorical
    dimensions (sex x environment x preference), balancing parsimony with AIC to select the most explanatory structure.

====================================================================================

#28  Cross-Tabulation
    file: 28_cross_age_answer.xlsx
  >> SCENARIO (narration):
    We examine age group and political leaning in a cross-tab. For the relationship
    of two qualitative variables, cross-tab is appropriate.
  >> VARIABLE SELECTION:
    - Variables: age
    - Variables: political

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi^2(6) = 9.29   p = 0.158 ns   Cramer's V = 0.115   H0 NOT REJECTED

>> COMMENTARY (narration):
    We examined two categorical variables (age group x political leaning) in a cross-tab: chi^2(6) = 9.29, p = 0.16,
    V = 0.12 -- no significant relationship, the variables are largely independent. Cross-tab + chi-square shows the
    direction and strength of the association between two qualitative variables at the cell level. In education/social
    research it is used to describe demography-attitude relationships; as here, "no relationship" is also a valuable
    finding (no need to treat the groups differently).

====================================================================================

#29  Multiple Response - Frequency
    file: 29_mr_frequency_applications.xlsx
  >> SCENARIO (narration):
    We analyze a multi-select activities question. For a multi-select question,
    multiple-response frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: activities (coklu yanit / multi-response)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)   multi-select "activities" question

>> COMMENTARY (narration):
    We analyzed a question where several options could be checked ("which activities do you take part in?"): each of 250
    participants gave one or more activities. Multiple-response frequency analysis gives the count of checks per option
    and both the response and case percentages separately (percentages sum to over 100, because one person picks
    multiple). In education/social surveys it is the standard method for correctly summarizing multi-select questions
    (activities joined, resources used).

====================================================================================

#30  Multiple Response - Cross-Tab
    file: 30_mr_categorical_region.xlsx
  >> SCENARIO (narration):
    We cross-tabulate the multi-response activities question by sex. To break a
    multi-select by a category, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: activities (coklu yanit / multi-response)
    - Grouping (categorical): sex

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: sex   |   multi-response "activities" x sex

>> COMMENTARY (narration):
    We cross-tabulated the multi-response activities question by a single category (sex): we compared the per-activity
    check rates for each group. A multiple-response cross-tab answers "does one group check certain options more often
    than another?". In education/social research it is used to compare groups' (sex, region) multi-select participation
    profiles (activities, resources).

====================================================================================

#31  Multiple Response x Multiple Response
    file: 31_mr_mr.xlsx
  >> SCENARIO (narration):
    We cross-tabulate two multi-response questions (interest x career) against each
    other. For many-to-many co-occurrence, this is appropriate.
  >> VARIABLE SELECTION:
    - Variables: interest (coklu / multi)
    - Variables: career (coklu / multi)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180   two multi-responses (interest x career) cross-tabulated

>> COMMENTARY (narration):
    We cross-tabulated two separate multi-response questions (interest areas x career goals) against each other. This is
    the most complex form of tabulation: on both axes a unit contributes to more than one cell. Multiple-response by
    multiple-response reveals many-to-many co-occurrences such as "which interest areas appear together with which
    career goals?". In education/guidance it is used to examine the matching pattern of nested preference bundles
    (interest x career).

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q_season.xlsx
  >> SCENARIO (narration):
    We test whether a binary outcome varies across courses in the same students. For
    3+ repeated binary measures, Cochran's Q is appropriate.
  >> VARIABLE SELECTION:
    - Columns: ders-bazli ikili sutunlar / course-wise binary columns

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000 ns   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the binary outcome (liked/not) measured for different courses in the same students varies
    significantly from course to course: Q = 0, p = 1.00 -- no difference across courses. Cochran's Q is the
    generalization of McNemar to more than two repeated conditions; it compares 3+ binary measures in the same unit. In
    education it is the right method for comparing whether the same students liked multiple courses/activities.

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5degisken.xlsx
  >> SCENARIO (narration):
    We examine all pairwise correlations among five continuous variables. For
    relationship direction and strength, the correlation matrix is appropriate.
  >> VARIABLE SELECTION:
    - Variables: motivation
    - Variables: study_hour
    - Variables: GPA
    - Variables: anxiety
    - Variables: satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables: motivation, study_hour, GPA, anxiety, satisfaction
    Strong relationship (|r| >= 0.75): NONE above this threshold

>> COMMENTARY (narration):
    We computed all pairwise Pearson correlations among five continuous variables (motivation, study hours, GPA,
    anxiety, satisfaction): none exceeded the 0.75 strong threshold -- the variables move largely independently. This
    too is a valuable finding: most measures carry different information, none substitutes for another. The correlation
    matrix summarizes the direction and strength of relationships at a glance; in education research it is the first
    step in spotting overlap among indicators (multicollinearity risk) and seeing which variables truly give separate
    signals.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Agreement
    file: 34_bland_altman_dbh.xlsx
  >> SCENARIO (narration):
    We examine the agreement of two tests (IQ_test_a/b) measuring the same
    intelligence. For test interchangeability, Bland-Altman is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: IQ_test_a
    - 2nd measure / group: IQ_test_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    IQ test A (IQ_test_a): mean = 100.52   |   IQ test B (IQ_test_b): mean = 99.20
    Bias (mean difference) ~ 1.32   |   limits of agreement (LoA) computed

>> COMMENTARY (narration):
    We examined how well two tests measuring the same intelligence (IQ_test_a vs IQ_test_b) agree: the mean systematic
    difference (bias) is ~1.3 points, and 95% limits of agreement were reported. Unlike correlation, Bland-Altman
    answers "can the two tests be used interchangeably?" -- high correlation does not mean agreement, there may be a
    systematic shift. In education/psychometrics it is the standard method for testing the interchangeability of two
    measuring instruments/tests.

====================================================================================

#35  Effect Size
    file: 35_effect_size_method.xlsx
  >> SCENARIO (narration):
    We measure the practical size of the score difference between two material
    groups. For importance beyond the p-value, effect size is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: score
    - Grouping (categorical): material

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.928 (large effect)   score difference between two material groups (material)

>> COMMENTARY (narration):
    We measured the practical size of the difference between two material groups, independent of the p-value, with
    Cohen's d: d = -0.93, a large effect. The p-value answers "is there a difference?"; effect size answers "how
    important is it?". Because large samples can produce significant but trivial differences, reporting effect size is
    essential. In education this measure clarifies whether the difference between two methods/materials is
    pedagogically noteworthy.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_physical_bio.xlsx
  >> SCENARIO (narration):
    We examine the joint structure between the cognitive-achievement set and the
    affective set. For the relationship between two multivariate sets, CCA is
    appropriate.
  >> VARIABLE SELECTION:
    - X variables: math
    - X variables: science
    - X variables: language
    - X variables: reading
    - X variables: logic
    - Y variables: motivation
    - Y variables: attitude
    - Y variables: satisfaction
    - Y variables: anxiety

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.985   chi^2(20) = 1407.32   p < .001 ***
    CC2: r = 0.965   chi^2(12) = 900.60    p < .001 ***
    Set-X (cognitive): math, science, language, reading, logic  |  Set-Y (affective): motivation, attitude, satisfaction, anxiety

>> COMMENTARY (narration):
    We resolved the joint structure between two multivariate measure sets -- cognitive achievement (math, science,
    language, reading, logic) and affective variables (motivation, attitude, satisfaction, anxiety) -- with canonical
    correlation. The first two canonical functions are very strong (r = 0.985 and 0.965, p < .001): the two sets are
    intensely related. CCA answers "how is one variable set related to another?" in a single step -- it is the
    multivariate-on-both-sides version of multiple regression. In education it reveals the latent relationship
    structure between the cognitive-achievement bundle and the affective bundle.

====================================================================================

#37  Correspondence Analysis
    file: 37_ca_plant_season.xlsx
  >> SCENARIO (narration):
    We map the relationship between noise level and occupation. To see the structure
    of two categorical variables, CA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: sound
    - Variables: occupation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.846   Dim1: 94.4%   Dim2: 5.6%   (2 dimensions)   sound x occupation

>> COMMENTARY (narration):
    We mapped the relationship in the cross-tab of two categorical variables (noise level x occupation) into a visual
    space with correspondence analysis: 94% of total inertia concentrates in one dimension -- the relationship can be
    summarized on essentially a single axis. CA positions the chi-square relationship in a two-dimensional space,
    showing which categories are close (co-occurring). In education/social research it is powerful for visually
    interpreting the structure of qualitative relationships such as group-occupation, attitude-region.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_18madde.xlsx
  >> SCENARIO (narration):
    We cluster eighteen items (self-efficacy/motivation/anxiety) by their
    similarity. To find the latent dimension structure, VarClus is appropriate.
  >> VARIABLE SELECTION:
    - Variables: self_efficacy_1..6, motivation_1..6, anxiety_1..6 (18 madde / items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: self_efficacy_1..6, motivation_1..6, anxiety_1..6)

>> COMMENTARY (narration):
    We clustered eighteen measurement items (self-efficacy, motivation, anxiety, 6 each) by how related they are: they
    grouped into 3 main dimensions -- most likely matching the natural self-efficacy/motivation/anxiety groups. VarClus
    groups the VARIABLES, not the observations -- by placing highly correlated items in the same cluster it reveals the
    latent dimensional structure of the data set. In education/psychometrics it is practical for reducing a long item
    set to a few core dimensions, and for spotting redundant items to shorten a scale.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_yield.xlsx
  >> SCENARIO (narration):
    We model GPA with four predictors (study, motivation, parent education,
    absenteeism). To explain a continuous outcome with multiple variables, multiple
    regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Predictor(s): study
    - Predictor(s): motivation
    - Predictor(s): parent_education
    - Predictor(s): absenteeism

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.616   Adj. R^2 = 0.608   predictors: study, motivation, parent_education, absenteeism

>> COMMENTARY (narration):
    We modeled GPA with four predictors (study, motivation, parent education, absenteeism) at once: the model explains
    61.6% of variance (Adj. R^2 = 0.61) -- strong explanatory power. Multiple regression gives each predictor's pure
    contribution to GPA with the others held constant; thus it answers "which factor really raises achievement?" while
    controlling confounders. In education it is the core method for identifying the drivers of achievement and for
    prediction.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_disease.xlsx
  >> SCENARIO (narration):
    We model university achievement (success/failure) with four predictors. For a
    binary outcome, logistic regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: university_achievement
    - Predictor(s): GPA
    - Predictor(s): study
    - Predictor(s): motivation
    - Predictor(s): sound

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 = 0.171   binary outcome: university_achievement   predictors: GPA, study, motivation, sound

>> COMMENTARY (narration):
    We modeled a binary outcome (university achievement: success/failure) with four predictors: the model has moderate
    explanatory power (pseudo R^2 = 0.17) and gives each predictor's effect on the odds. Logistic regression replaces
    linear regression when the outcome is binary; coefficients are converted to Odds Ratios to read "how many times
    does the success odds change per unit increase in this variable?". In education it is the core model for predicting
    yes/no outcomes such as passing, admission, dropout.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Count (Poisson) Regression
    file: 41_poisson_insect.xlsx
  >> SCENARIO (narration):
    We model the disciplinary-event count with age and motivation. For a count
    outcome, Poisson regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: discipline_count
    - Predictor(s): age
    - Predictor(s): motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 180.91   Deviance = 113.68   count outcome: discipline_count   predictors: age, motivation

>> COMMENTARY (narration):
    We modeled a count variable (disciplinary-event count) with age and motivation. Poisson regression is the right
    model when the outcome is a count (0,1,2,... items); linear regression is unsuitable because it can produce
    negative/fractional predictions. Coefficients give the effect on the count rate. In education it is used to explain
    count outcomes such as disciplinary events, absence count, application count.

====================================================================================

#42  Multinomial Logistic
    file: 42_multinomial_type.xlsx
  >> SCENARIO (narration):
    We model the multi-category occupational preference with two predictors. For a
    nominal multi-class outcome, multinomial logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: occupation
    - Predictor(s): math
    - Predictor(s): art

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 252.96   Reference class: 'artist'   outcome: occupation (>2 categories)   predictors: math, art

>> COMMENTARY (narration):
    We modeled a more-than-two-category outcome (occupational preference) with two predictors (math, art). Multinomial
    logistic compares each category against a reference class (here 'artist') with a separate logistic equation;
    coefficients are read as "as X increases, how does the chance of being in this occupation change relative to the
    reference?". When the outcome is nominal with more than two classes (occupation, preference type) it is the right
    choice. In education/guidance it is the standard model for multi-option class prediction.

====================================================================================

#43  Ordinal Logistic
    file: 43_ordinal_quality.xlsx
  >> SCENARIO (narration):
    We model the ordinal motivation level with teacher support and school
    opportunity. For an ordinal outcome, ordinal logistic is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: motivation_level
    - Predictor(s): teacher_support
    - Predictor(s): school_opportunity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 466.81   ordinal outcome: motivation_level   predictors: teacher_support, school_opportunity

>> COMMENTARY (narration):
    We modeled an ordinal outcome (motivation level: low/medium/high) with teacher support and school opportunity.
    Ordinal logistic uses the ORDER information between categories (which multinomial ignores); with a "proportional
    odds" assumption it explains all thresholds with one coefficient set. In education it is the right and more
    powerful choice for modeling naturally ordered outcomes such as motivation level, achievement class, grade.

====================================================================================

#44  PLS Regression
    file: 44_pls_spectral_NDVI.xlsx
  >> SCENARIO (narration):
    We predict achievement from 12 correlated behavior items. For many highly
    correlated predictors, PLS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: achievement
    - Predictor(s): behavior_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 (train) = 0.877   R^2 (5-fold CV) = 0.860   predictors: behavior_01..12 (12 behavior items)

>> COMMENTARY (narration):
    We predicted achievement from 12 correlated behavior items. Because the items are highly correlated
    (multicollinearity), classic regression becomes unstable; PLS reduces them to a few latent components and regresses
    on those. With cross-validated R^2 = 0.86, the model is both strong and generalizable. PLS is ideal when predictors
    are numerous or highly correlated; in education/psychometrics it is widely used for predicting an outcome from many
    correlated items.

====================================================================================

#45  Probit Regression
    file: 45_probit_dose_response.xlsx
  >> SCENARIO (narration):
    We estimate the probability of passing from study hours with a probit model. For
    a binary outcome, probit (normal link) is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: passed
    - Predictor(s): study_hour

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 97.53   Pseudo R^2 (McFadden) = 0.603   binary outcome: passed   predictor: study_hour

>> COMMENTARY (narration):
    We estimated the probability of passing (passed: yes/no) from study hours with a probit model: the model is very
    strong (pseudo R^2 = 0.60). Probit, like logistic, applies to binary outcomes; the difference is that its link
    function is the normal distribution. Results are usually similar; probit is preferred where an underlying latent
    normal variable is natural. In education it is a robust alternative to logistic for binary decision/outcome
    modeling.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_investment.xlsx
  >> SCENARIO (narration):
    We model a zero-piled education-spending variable with income and child count.
    For a censored outcome, tobit is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: education_tuition
    - Predictor(s): income
    - Predictor(s): child_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 475.35   Left censor   outcome: education_tuition (censored)   predictors: income, child_count

>> COMMENTARY (narration):
    We modeled a lower-bounded (zero-piled) education-spending variable with income and child count. Tobit is for
    "censored" dependent variables that pile up at a threshold (here 0); ordinary regression gives biased estimates by
    ignoring this pile-up. In education/social research it is the right model for floor-effect outcomes such as zero
    spending; coefficients reflect the true (uncensored) relationship.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_NPK.xlsx
  >> SCENARIO (narration):
    We model the achievement score with study and age in a Bayesian framework. To
    express uncertainty probabilistically, Bayesian regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Predictor(s): study
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2 posterior mean = 16.95 (sd = 2.42)   outcome: score   predictors: study, age
    coefficient posteriors + P(beta>0) reported

>> COMMENTARY (narration):
    We modeled the achievement score with study and age in a Bayesian framework: instead of point estimates we obtained
    each coefficient's full posterior distribution and the "probability the effect is positive". The Bayesian approach
    expresses uncertainty directly in probability language and can incorporate prior knowledge. In education, when the
    sample is small or prior-term information is valuable, it offers intuitive interpretations like "the effect is
    probably positive".

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear_growth.xlsx
  >> SCENARIO (narration):
    We model the S-shaped relationship of word count with age via a logistic growth
    curve. For a curved relationship, nonlinear regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: word_count
    - Predictor(s): age
    - Value: fonksiyon/function: logistic

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Function: y = K / (1 + exp(-r*(x-x0)))  (logistic growth)   R^2 = 0.966

>> COMMENTARY (narration):
    We modeled the relationship of an outcome (word_count) with age not as a straight line but as an S-shaped logistic
    growth curve: the fit is very high (R^2 = 0.97). Nonlinear regression fits a theoretical function form directly to
    the data when the relationship is curved (saturation, threshold, exponential growth) and makes the parameters
    (ceiling K, rate r, inflection x0) interpretable. In education/development it is the right tool for modeling
    S-shaped processes such as language acquisition and learning curves.

====================================================================================

#49  Ridge Regression
    file: 49_ridge_soil_yield.xlsx
  >> SCENARIO (narration):
    We predict motivation from 15 correlated items with ridge. For
    multicollinearity, ridge is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: motivation
    - Predictor(s): item_p01..15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 1.0   R^2 = 0.617   predictors: item_p01..15 (15 items)

>> COMMENTARY (narration):
    We predicted motivation from 15 correlated items with ridge regression (R^2 = 0.62). Ridge adds an L2 penalty to
    shrink all coefficients in magnitude without zeroing them; this prevents the instability caused by high correlation
    among predictors (multicollinearity). It gives more stable and generalizable estimates where classic regression's
    coefficients balloon and flip sign. In education/psychometrics it is preferred for prediction with many co-varying
    items.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso.xlsx
  >> SCENARIO (narration):
    We model achievement from 30 candidate features with lasso, selecting the
    important ones. For automatic variable selection, lasso is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: achievement
    - Predictor(s): x01..30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.1   R^2 = 0.726   18 of 30 features shrunk to zero (automatic variable selection)

>> COMMENTARY (narration):
    We modeled achievement from 30 candidate features with lasso regression: R^2 = 0.73, and lasso shrank 18 of the 30
    feature coefficients exactly to zero, selecting only 12 effective variables. This is its difference from ridge:
    because lasso can zero coefficients, it performs prediction and variable selection at the same time. In
    education/social research it is very useful for automatically winnowing the "few truly important factors" out of
    many candidate indicators (a sparse, interpretable model).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_nitrogen_yield.xlsx
  >> SCENARIO (narration):
    We test the parent support -> motivation -> achievement chain. To resolve the
    intermediate mechanism, mediation analysis is appropriate.
  >> VARIABLE SELECTION:
    - Predictor(s): parent_support
    - Value: M (araci/mediator): motivation
    - Dependent variable: achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 5.64   95% CI [3.91, 7.57]   (excludes zero -> significant)
    X: parent_support  M: motivation  Y: achievement

>> COMMENTARY (narration):
    We tested the chain "parent support (X) -> motivation (M) -> achievement (Y)": the indirect effect is 5.64, its 95%
    CI excludes zero -- so support's effect on achievement occurs substantially through motivation. Mediation analysis
    resolves "why/how does X affect Y?" through an intermediate mechanism. In education/psychology it is powerful for
    understanding which intermediate process (motivation, self-efficacy) a support's/intervention's effect flows
    through.

====================================================================================

#52  Path Analysis
    file: 52_path_climate_yield.xlsx
  >> SCENARIO (narration):
    We test direct/indirect relationships as a single causal diagram. For a
    relationship network, path analysis is appropriate.
  >> VARIABLE SELECTION:
    - Value: achievement ~ self_efficacy + motivation + effort
    - Value: motivation ~ self_efficacy

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.866   RMSEA = 0.243   model: achievement ~ self_efficacy + motivation + effort; motivation ~ self_efficacy

>> COMMENTARY (narration):
    We tested the direct and indirect relationships among several variables as a single causal diagram. The fit indices
    (CFI = 0.87, RMSEA = 0.24) show the model fits the data partially, with room for improvement. Path analysis
    estimates the whole relationship network at once instead of separate regressions; it shows how variables affect
    each other and a common outcome. In education/psychology it is used to test theory-based relationship models
    (self-efficacy -> motivation -> achievement).

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_plot_year.xlsx
  >> SCENARIO (narration):
    We model GPA measured repeatedly across terms in the same students, taking
    student as a random effect. For repeated/nested data, LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Predictor(s): term
    - Cluster: student_id (random)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC = 0.943   Group variance (random intercept) = 0.247   outcome: GPA   fixed: term   group: student_id

>> COMMENTARY (narration):
    We modeled GPA measured repeatedly across terms in the same students, taking student identity as a random effect.
    ICC = 0.94 is very high: almost all GPA variability comes from between-student differences, while within-student
    terms are very similar. LMM correctly handles dependency in nested/repeated (measurements within student) data; it
    solves the "independence" assumption that ordinary regression violates via random effects. In education it is the
    right choice for panel/repeated-measure data.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation.xlsx
  >> SCENARIO (narration):
    Instead of deleting missing data we fill it with 5 plausible value sets. To
    handle missingness without bias, multiple imputation is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Predictor(s): study
    - Predictor(s): motivation
    - Predictor(s): anxiety
    - Predictor(s): parent_education

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (imputation) = 5   outcome: GPA   predictors: study, motivation, anxiety, parent_education

>> COMMENTARY (narration):
    Instead of deleting missing data, we analyzed by producing 5 plausible value sets (m = 5) and combining the
    results. Multiple imputation -- unlike filling gaps with a single estimate (which ignores uncertainty) -- accounts
    for imputation uncertainty too, yielding unbiased estimates and correct standard errors. In education/social survey
    data it is the modern standard for handling the inevitable gaps without shrinking the sample or distorting results.

====================================================================================

#55  GEE
    file: 55_gee_drug_yield.xlsx
  >> SCENARIO (narration):
    We model motivation in repeated visit measures of the same students. For a
    population-average effect, GEE is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: motivation
    - Predictor(s): visit
    - Cluster: student_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coef = 0.032   p = 0.042 *   QIC = 305.56   outcome: motivation   group: student_id

>> COMMENTARY (narration):
    We modeled motivation in repeated visit measures of the same students; the visit effect is significant (b = 0.032,
    p = 0.042). GEE estimates the POPULATION-AVERAGE effect rather than individual effects in repeated/clustered data
    and corrects within-group correlation with a "working correlation structure". While LMM focuses on individual
    random effects, GEE focuses on the average trend. In education it is preferred for population-level questions like
    "what is the time/visit effect in the average student?".

====================================================================================

#56  GLMM
    file: 56_glmm_insect.xlsx
  >> SCENARIO (narration):
    We model a repeatedly measured discipline count in the same students with month
    and intervention, taking student as a random effect. For repeated counts, GLMM
    is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: discipline
    - Predictor(s): month
    - Predictor(s): intervention
    - Cluster: student_id (Poisson)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    intervention coef = -0.206   p = 0.077 ns   outcome: discipline (count, Poisson)   group: student_id

>> COMMENTARY (narration):
    We modeled a COUNT outcome (disciplinary-event count) measured repeatedly in the same students with month and
    intervention, taking student as a random effect; the intervention effect is borderline but non-significant
    (p = 0.077). GLMM extends LMM to non-normal outcomes (count, binary): it handles both the distribution (Poisson)
    and the clustering (random effect) at the same time. In education it is the right model for repeatedly measured
    binary/count outcomes (monthly discipline count, dropout).

====================================================================================

#57  Elastic Net
    file: 57_elasticnet.xlsx
  >> SCENARIO (narration):
    We model an index from 40 predictors with Elastic Net. For many clustered
    predictors, Elastic Net is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: index
    - Predictor(s): x01..40

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.075   L1 ratio = 0.50   R^2 (train) = 0.604   40 predictors

>> COMMENTARY (narration):
    We modeled an index from 40 predictors with Elastic Net. Elastic Net blends the ridge (L2) and lasso (L1) penalties
    (L1 ratio = 0.5): it both keeps groups of correlated variables together (ridge property) and zeroes out redundant
    ones (lasso property). Cross-validated generalizability was measured. In education/social research it is a balanced
    choice when there are many predictors clustered among themselves, where lasso or ridge alone is insufficient.

====================================================================================

#58  Robust Regression
    file: 58_robust_area_yield.xlsx
  >> SCENARIO (narration):
    We model the achievement score with study, down-weighting outliers. For data
    with outliers, robust regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Predictor(s): study

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust Intercept = 51.24   OLS Intercept = 49.14   outcome: score   predictor: study

>> COMMENTARY (narration):
    When modeling the achievement score with study, we used robust regression to prevent outliers from distorting the
    estimate. The gap between the robust and OLS intercepts (51.24 vs 49.14) shows a few outlying observations pull the
    classic estimate; the robust method down-weights them to reflect the "typical" relationship. In education/social
    research it is the right way to get robust estimates without deleting outliers (abnormal record) in data that
    contains them.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_yield.xlsx
  >> SCENARIO (narration):
    We model different points of GPA's distribution (lower 10%, median, upper 90%)
    separately. For varying effects, quantile regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Predictor(s): study
    - Predictor(s): motivation
    - Quantiles: 0.10 / 0.50 / 0.90

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    separate slopes reported for q=0.10, q=0.50, q=0.90   outcome: GPA   predictors: study, motivation

>> COMMENTARY (narration):
    We modeled not just GPA's mean but different points of the distribution (lower 10%, median, upper 90%) separately.
    Predictor effects can vary by quantile -- a factor may be strong for low-achieving students and weak for
    high-achieving ones. Quantile regression gives the true picture when the "mean effect" is misleading (effect varies
    across the distribution). In education it shows what classic regression misses by examining inequality and
    extreme-group behavior.

====================================================================================

#60  ROC Curve
    file: 60_roc_quality.xlsx
  >> SCENARIO (narration):
    We assess how well a continuous score separates a binary outcome with ROC. For
    discrimination and threshold selection, ROC is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: successful
    - Predictor(s): composite

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.921 (excellent discrimination)   Youden optimum threshold = 0.710   Sensitivity = 0.787, 1-Specificity = 0.080
    n = 250 (positive 150, negative 100)

>> COMMENTARY (narration):
    We assessed how well a continuous score (composite) separates a binary outcome (successful) with a ROC curve:
    AUC = 0.92, excellent discrimination; the optimum decision threshold was set at 0.71 via Youden. ROC shows the
    sensitivity-specificity trade-off at all possible thresholds and evaluates the model without being tied to a single
    threshold. In education it is the standard tool for measuring a risk/achievement model's discriminative power and
    selecting the best decision threshold.

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss_type_distribution.xlsx
  >> SCENARIO (narration):
    We measure a classification model's discrimination with TSS. For a fair measure
    under class imbalance, TSS is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: pass
    - Predictor(s): GPA
    - Predictor(s): absenteeism

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.33 (acceptable)   Sensitivity = 0.73   Specificity = 0.60   N = 200
    binary outcome: pass   predictors: GPA, absenteeism

>> COMMENTARY (narration):
    We measured a classification model's (pass/fail) discrimination with TSS: TSS = 0.33 (sensitivity 0.73, specificity
    0.60). TSS = sensitivity + specificity - 1; it excludes chance-expected success and gives a fair performance
    measure even with imbalanced classes. While accuracy can be biased (it anchors to the majority class), TSS
    evaluates the power to capture both positives and negatives together. In education it is robust for reporting the
    true discriminative power of risk/achievement classification models.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity.xlsx
  >> SCENARIO (narration):
    We evaluate a classifier's predictions against true labels. For detailed
    performance, the confusion matrix is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: actual_label
    - 2nd measure / group: prediction_label

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.93   F1 = 0.918   (true label vs predicted label)

>> COMMENTARY (narration):
    We evaluated a classifier's predictions against true labels via a confusion matrix: accuracy 93%, F1 = 0.92. The
    confusion matrix gathers true/false positives and negatives in one table, from which sensitivity, specificity,
    precision and F1 are derived. Because a single accuracy number can mislead (especially with imbalanced classes),
    this metric set shows where the model errs. In education it is fundamental for detailed reporting of
    prediction/classification model performance.

====================================================================================

#63  Random Forest
    file: 63_rf_tree_type.xlsx
  >> SCENARIO (narration):
    We classify occupational preference from five features with a random forest. For
    nonlinear prediction, random forest is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: occupation
    - Predictor(s): math
    - Predictor(s): language
    - Predictor(s): GPA
    - Predictor(s): reading
    - Predictor(s): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.757   outcome: occupation   predictors: math, language, GPA, reading, age

>> COMMENTARY (narration):
    We classified occupational preference from five features with a random forest: accuracy 75.7%. Random forest
    combines the votes of hundreds of decision trees; it automatically captures nonlinear relationships and
    interactions and also gives a variable-importance ranking. It reduces a single tree's overfitting by averaging. In
    education/guidance it is widely used for complex, nonlinear prediction problems (occupation/orientation prediction)
    because it offers both high accuracy and "which variable matters" information.

====================================================================================

#64  Support Vector Machines (SVM)
    file: 64_svm_health.xlsx
  >> SCENARIO (narration):
    We classify university-admitted/not from four features with SVM. For
    well-separable classes, SVM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: university_passed
    - Predictor(s): GPA
    - Predictor(s): math
    - Predictor(s): trial
    - Predictor(s): study

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: university_passed   predictors: GPA, math, trial, study

>> COMMENTARY (narration):
    We classified university-admitted/not from four features with SVM: accuracy 100% -- the classes are perfectly
    separable with these features (very high accuracy should be confirmed against overfitting via cross-validation).
    SVM finds the decision boundary separating classes with the widest margin; with the kernel trick it can also do
    nonlinear separation. In education it is a strong, stable classifier for well-separable class problems, especially
    on small-to-medium data.

====================================================================================

#65  Gradient Boosting
    file: 65_gradient_boosting.xlsx
  >> SCENARIO (narration):
    We classify dropout risk from four predictors with gradient boosting. For top
    accuracy, gradient boosting is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: dropout_risk
    - Predictor(s): GPA
    - Predictor(s): absenteeism
    - Predictor(s): sound
    - Predictor(s): parent_joint

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.982   outcome: dropout_risk   predictors: GPA, absenteeism, sound, parent_joint

>> COMMENTARY (narration):
    We classified dropout risk (dropout_risk) from four predictors with gradient boosting: accuracy 98.2%. Gradient
    boosting adds weak trees sequentially -- each new tree corrects the previous model's errors; this is why it wins
    most prediction competitions. While random forest votes in parallel, boosting reduces error step by step. In
    education it is preferred where the highest predictive accuracy is sought (dropout-risk scoring, achievement
    prediction); it requires tuning against overfitting.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_4kume.xlsx
  >> SCENARIO (narration):
    We cluster students by four features. For student segmentation, k-means is
    appropriate.
  >> VARIABLE SELECTION:
    - Variables: GPA
    - Variables: study
    - Variables: math
    - Variables: trial
    - Number of clusters: 4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.459   variables: GPA, study, math, trial

>> COMMENTARY (narration):
    We split students into 4 clusters by four features (GPA, study, math, trial): silhouette = 0.46, reasonable-
    moderate separation. K-means assigns observations to the nearest cluster center and iteratively updates the
    centers; it groups similar students into natural clusters. No labels are needed (unsupervised). In education it is
    the core method for student segmentation and discovering learning profiles; silhouette checks the appropriateness
    of the cluster count.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5tur.xlsx
  >> SCENARIO (narration):
    We hierarchically cluster students by ten behavior items. For nested group
    structure, hierarchical clustering is appropriate.
  >> VARIABLE SELECTION:
    - Variables: behavior_01..10
    - Number of clusters: 5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 5   Silhouette = 0.123   variables: behavior_01..10 (10 behavior items)

>> COMMENTARY (narration):
    We split students into 5 groups by ten behavior items with hierarchical clustering (silhouette = 0.12, weak
    separation -- groups partly overlap). Hierarchical clustering merges observations step by step to build a tree
    (dendrogram); its difference from k-means is not having to fix the cluster count in advance and seeing the nested
    structure. In education it is used to explore a hierarchy of student groups (main group -> sub-group) and to read
    the natural cluster count from the dendrogram.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan.xlsx
  >> SCENARIO (narration):
    We cluster school locations with density-based DBSCAN. For spatial clusters and
    outliers, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Variables: lat
    - Variables: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n_clusters = 4   Silhouette = 0.916   variables: lat, lon (school location)

>> COMMENTARY (narration):
    We clustered school locations (lat, lon) with density-based DBSCAN: 4 dense clusters, silhouette = 0.92, very good
    separation. Unlike k-means, DBSCAN does not require the cluster count in advance, can find clusters of any shape,
    and marks sparse points as "noise". In education/social geography it is ideal for detecting spatial concentrations
    (school clusters, access zones) and isolating outlier locations.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca_6ozellik.xlsx
  >> SCENARIO (narration):
    We reduce six correlated achievement indicators to a few components with PCA.
    For dimensionality reduction, PCA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: math
    - Variables: science
    - Variables: language
    - Variables: reading
    - Variables: logic
    - Variables: art

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 explains 88.3% of variance   variables: math, science, language, reading, logic, art

>> COMMENTARY (narration):
    We reduced six correlated achievement indicators to a few components with PCA: the first component alone explains
    88% of variance -- so these six course scores largely reflect a single latent dimension (general academic ability).
    PCA transforms correlated variables into mutually independent components; it reduces dimensions, eases
    visualization and resolves multicollinearity. In education it is fundamental for summarizing multi-course
    achievement and building a "general achievement index".

====================================================================================

#70  t-SNE
    file: 70_tsne_5tur.xlsx
  >> SCENARIO (narration):
    We reduce 12-dimensional data to 2 dimensions with t-SNE for visualization. For
    nonlinear cluster discovery, t-SNE is appropriate.
  >> VARIABLE SELECTION:
    - Variables: item_01..12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence = 0.544 (good)   high-dimensional 12 items embedded into 2 dimensions

>> COMMENTARY (narration):
    We reduced 12-dimensional data to 2 dimensions for visualization with t-SNE (KL = 0.54, good quality). t-SNE tries
    to preserve high-dimensional neighborhoods, placing similar observations near and dissimilar ones far; it reveals
    nonlinear cluster structures visually that PCA misses. Interpretation is visual (the axes have no absolute
    meaning). In education/psychometrics it is used for exploratory visualization of hidden student clusters in
    high-dimensional scale/behavior data.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds_3grup.xlsx
  >> SCENARIO (narration):
    We place the inter-student similarity structure onto a 2D map. For similarity
    maps, MDS is appropriate.
  >> VARIABLE SELECTION:
    - Variables: GPA
    - Variables: study
    - Variables: math
    - Variables: science
    - Variables: language

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.052 (acceptable)   variables: GPA, study, math, science, language

>> COMMENTARY (narration):
    We placed the distance/similarity structure among students onto a 2-dimensional map with low stress (0.05): low
    stress means the map represents the true distances well. MDS positions observations in an interpretable space by
    preserving inter-observation distances; its difference from t-SNE is the aim of preserving global distance
    structure. In education it is used to draw student/group similarity maps and for positioning between groups.

====================================================================================

#72  UMAP
    file: 72_umap_5tip.xlsx
  >> SCENARIO (narration):
    We reduce 25-dimensional behavior data to 2 dimensions with UMAP. For both local
    and global structure, UMAP is appropriate.
  >> VARIABLE SELECTION:
    - Variables: behavior_01..25

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    25 behavior items embedded into 2 dimensions   (local-global balance via n_neighbors / min_dist)

>> COMMENTARY (narration):
    We reduced 25-dimensional behavior data to 2 dimensions with UMAP. UMAP does nonlinear dimensionality reduction
    like t-SNE but preserves both local and global structure better and is faster. Larger n_neighbors emphasizes
    broader groups, larger min_dist emphasizes the gaps between clusters. In education/psychometrics it is a modern
    choice for visualizing high-dimensional behavior/scale data and exploring natural cluster structure.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_21.xlsx
  >> SCENARIO (narration):
    We measure the internal consistency of a 21-item scale. For scale reliability,
    Cronbach's alpha is appropriate.
  >> VARIABLE SELECTION:
    - Variables: item_01..21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach alpha = 0.972 (excellent internal consistency)   21 items

>> COMMENTARY (narration):
    We measured the internal consistency of a 21-item scale with Cronbach's alpha: alpha = 0.97, excellent -- the items
    measure the same construct consistently. Cronbach's alpha shows how much the items of a scale "move together" (their
    equivalence); above 0.70 is considered acceptable. (A very high value can also signal item redundancy.) In
    education/psychometrics it is the standard index for reporting the reliability of attitude/engagement scales.

====================================================================================

#74  Likert Scale Analysis
    file: 74_likert_3boyut.xlsx
  >> SCENARIO (narration):
    We analyze a 15-item Likert set of three sub-scales
    (motivation/self-efficacy/anxiety). For ordinal scale summary and reliability,
    Likert analysis is appropriate.
  >> VARIABLE SELECTION:
    - Variables: motivation_1..5 / self_efficacy_1..5 / anxiety_1..5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of items k = 15   Cronbach alpha = 0.769 (acceptable)

>> COMMENTARY (narration):
    We analyzed a 15-item Likert set made of three sub-scales (motivation, self-efficacy, anxiety): reliability is
    acceptable (alpha = 0.77), and item distributions and central tendencies were reported. Likert analysis describes
    ordinal scale responses (1-5) with appropriate summaries (median, distribution, pile-up) and checks scale
    reliability. In education/psychometrics it is used for correctly summarizing and interpreting attitude/perception
    surveys (motivation, anxiety).

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18.xlsx
  >> SCENARIO (narration):
    We discover the latent factors behind eighteen items. For scale-structure
    discovery, EFA is appropriate.
  >> VARIABLE SELECTION:
    - Variables: q01..18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO = 0.908 (excellent)   factor analysis appropriate   18 items (q01..18)

>> COMMENTARY (narration):
    We applied EFA to discover how many latent factors lie behind eighteen items: KMO = 0.91, the data is very suitable
    for factor analysis. EFA reduces observed items to a few unobserved "factors"; it reveals which items measure the
    same dimension. In education/psychometrics it is fundamental when developing a new scale (discovering the survey's
    structure) and mapping item groups to theoretical dimensions; KMO and Bartlett are prerequisite tests.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3uzman.xlsx
  >> SCENARIO (narration):
    We measure the consistency of scores from three teachers. For inter-rater
    reliability on continuous scores, ICC is appropriate.
  >> VARIABLE SELECTION:
    - Variables: teacher_a
    - Variables: teacher_b
    - Variables: teacher_c

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.961 (Excellent)   95% CI [0.940, 0.980]   3 raters (teacher_a/b/c)

>> COMMENTARY (narration):
    We measured the consistency of scores three teachers gave to the same students with ICC: ICC = 0.96, excellent --
    the teachers score almost identically. Unlike kappa, ICC measures inter-rater reliability on CONTINUOUS scores and
    can assess both consistency and absolute agreement. In education it is the standard index for determining how
    reliable/interchangeable different teachers' scores (achievement level, performance class) are.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_12_3.xlsx
  >> SCENARIO (narration):
    We test whether a predefined three-factor structure (self/academic/social
    self-concept) fits the data. For construct validity, CFA is appropriate.
  >> VARIABLE SELECTION:
    - Value: self: self_1..4
    - Value: academic: academic_1..4
    - Value: social: social_1..4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 1.001   RMSEA = 0.000   3 factors: self / academic / social (4 items each)

>> COMMENTARY (narration):
    Unlike EFA, we TESTED whether a pre-defined three-factor structure (self/academic/social self-concept) fits the
    data with CFA: the fit indices are excellent (CFI = 1.00, RMSEA = 0.00). CFA tests a theoretical scale model --
    which item loads on which factor is fixed in advance, and the question is "does the model fit the data?". In
    education/psychometrics it is a mandatory step for confirming the construct validity of a developed scale (do the
    self-concept dimensions match theory).

====================================================================================

#78  Survey Mean
    file: 78_survey_means.xlsx
  >> SCENARIO (narration):
    We estimate the mean PISA reading score in a stratified/weighted survey. For a
    complex sample, design-based mean is appropriate.
  >> VARIABLE SELECTION:
    - Variables: PISA_reading
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PISA_reading: M_hat = 458.99   SE = 3.99   95% CI [451.14, 466.83]   CV 0.87%   weight: weight, stratum: region

>> COMMENTARY (narration):
    In a stratified/weighted survey design we estimated the mean PISA reading score while accounting for the design:
    M = 458.99, with a Taylor-linearization SE and 95% CI [451, 467]. Complex-sample methods account for unequal
    selection probabilities (weights) and stratification; ignoring these biases the standard errors. In education it is
    the right way to produce correct point estimates and confidence intervals from national/regional representative
    surveys (like PISA, TIMSS).

====================================================================================

#79  Survey Frequency
    file: 79_survey_freq.xlsx
  >> SCENARIO (narration):
    We estimate university-intention category proportions accounting for the design.
    For a complex sample, design-based frequency is appropriate.
  >> VARIABLE SELECTION:
    - Variables: university_intention
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    weighted proportions + Taylor SE + 95% CI for university_intention categories
    (e.g. 'maybe' p_hat = 0.363, 'certain' p_hat = 0.095)

>> COMMENTARY (narration):
    We estimated the category proportions of a categorical survey question (university intention) under sampling
    weights and stratification; each proportion has a design-based SE and confidence interval. Complex-survey frequency
    analysis, unlike a simple percentage, estimates population proportions without bias by accounting for the sampling
    design. In education it is used to correctly report intention/preference distributions from representative surveys.

====================================================================================

#80  Survey Total
    file: 80_survey_total.xlsx
  >> SCENARIO (narration):
    We estimate the population total (total teachers) from the sample. For a complex
    sample, design-based total is appropriate.
  >> VARIABLE SELECTION:
    - Variables: teacher_count
    - Weight: weight
    - Stratum: stratum

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    teacher_count: T_hat = 67,866   SE = 1,314.75   95% CI [65,278, 70,454]   stratum: stratum

>> COMMENTARY (narration):
    We estimated the population TOTAL (total teacher count) from the sample: T = 67,866, 95% CI [65,278, 70,454].
    Weights tell how many population units each observation represents; the total is estimated by summing those
    weights, with uncertainty reported via Taylor SE. In education it is the right method for producing
    population-scaled totals (total teachers, total students) from a sample.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg.xlsx
  >> SCENARIO (narration):
    We regress the PISA score accounting for the design. For relationships in
    complex samples, design-based regression is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: PISA
    - Predictor(s): age
    - Predictor(s): sound
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.185   age coef = 14.12 (p < .001)   outcome: PISA   predictors: age, sound   weight: weight

>> COMMENTARY (narration):
    We regressed a survey outcome (PISA score) while accounting for the sampling design (weight, stratum): age is a
    significant predictor (b = 14.12, p < .001), and the model explains 18.5% of variance. Design-based regression
    incorporates weights and the cluster/stratum structure into coefficient and standard-error calculation; ordinary
    regression ignores these and gives biased inference. In education it is the right way to model relationships in
    representative survey data.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic.xlsx
  >> SCENARIO (narration):
    We model university admission with design-weighted logistic. For binary outcomes
    in complex samples, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: university_passed
    - Predictor(s): GPA
    - Predictor(s): sound
    - Weight: weight
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    GPA: OR = 5.58   p < .001 ***   binary outcome: university_passed   predictors: GPA, sound   weight: weight

>> COMMENTARY (narration):
    We modeled a binary survey outcome (university admitted) with design-weighted logistic regression: each one-unit
    increase in GPA raises the admission odds ~5.6-fold (OR = 5.58, p < .001) -- a strong predictor. This method
    extends logistic regression to complex sample designs -- weight and stratum are reflected in the standard errors.
    In education it is used to correctly estimate the probability of a binary outcome (admitted/not) from
    representative surveys.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_temperature_yield.xlsx
  >> SCENARIO (narration):
    We model the achievement score with study via a flexible curve. For a nonlinear
    effect, GAM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Predictor(s): study (smooth)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R^2 (explained) = 0.756   outcome: score   predictor: study (smooth term)

>> COMMENTARY (narration):
    We modeled the achievement score with study, without assuming a straight line in advance, via a flexible curve
    (smooth): the explained variance is high (pseudo R^2 = 0.76). GAM extends linear regression -- it models each
    predictor's effect as a smooth function learned from the data, capturing curved relationships without losing
    interpretability. In education it is a more explanatory choice than black-box models for relationships where the
    effect is nonlinear (saturation, threshold, optimum study).

====================================================================================

#84  Discriminant Analysis
    file: 84_diskriminant_3sinif.xlsx
  >> SCENARIO (narration):
    We classify student class from five continuous measures. To assign to predefined
    groups, discriminant analysis is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: class
    - Predictor(s): math
    - Predictor(s): science
    - Predictor(s): GPA
    - Predictor(s): absenteeism
    - Predictor(s): study

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.00   outcome: class   predictors: math, science, GPA, absenteeism, study

>> COMMENTARY (narration):
    We classified student class from five continuous measures with discriminant analysis: accuracy 100% -- the classes
    are fully separable with these measures (overfitting risk should be checked via cross-validation). Discriminant
    analysis finds the linear combinations that best separate groups; it both classifies and shows which variable is
    most influential in separation. In education it is used to assign new students to predefined groups and to identify
    discriminating features.

====================================================================================

#85  Conditional Logit
    file: 85_conditional_logit.xlsx
  >> SCENARIO (narration):
    We model students' choices among university alternatives by option features. For
    discrete-choice data, conditional logit is appropriate.
  >> VARIABLE SELECTION:
    - Chooser: student_id
    - Choice: chosen
    - Alternative features: fee
    - Alternative features: distance_km

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McFadden pseudo R^2 = 1.00   chooser: student_id   choice: chosen   features: fee, distance_km

>> COMMENTARY (narration):
    We modeled the choices students made among university alternatives by the alternatives' features (fee, distance)
    with a conditional logit. This model is for "discrete choice" data where each individual picks one from a choice
    set; it estimates how an option's features affect its probability of being chosen. Its difference from standard
    logistic is that the choice is conditional on the individual's option set. In education/social research it is the
    core method for school/program preference modeling (fee-distance effect).

====================================================================================

#86  Kaplan-Meier Survival
    file: 86_km_tree.xlsx
  >> SCENARIO (narration):
    We examine students' time to an event and group differences. For censored time
    data, Kaplan-Meier is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_term
    - Event: graduate
    - Grouping (categorical): department

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 8.4 terms   Log-rank chi^2(2) = 1.17, p = 0.557 ns   time: duration_term, event: graduate, group: department

>> COMMENTARY (narration):
    We examined the time students take to reach an event (graduate) with Kaplan-Meier: median survival 8.4 terms;
    department survival curves were compared with the log-rank test (p = 0.56, no difference). KM correctly handles
    censored time data (those whose event has not yet occurred); a plain average ignores these observations and is
    biased. In education it is fundamental for time-to-graduation, dropout-timing and retention analysis.

====================================================================================

#87  Cox Proportional Hazards
    file: 87_cox_hazard.xlsx
  >> SCENARIO (narration):
    We examine dropout risk with continuous and categorical predictors in a Cox
    model. For multiple predictors in censored time, Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_term
    - Event: dropout
    - Predictor(s): age
    - Predictor(s): GPA
    - Predictor(s): scholarship
    - Predictor(s): family_support

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.649   GPA: HR = 0.569   p = 0.002 **   (higher GPA ~ lower dropout risk)
    time: duration_term, event: dropout   predictors: age, GPA, scholarship, family_support

>> COMMENTARY (narration):
    We examined dropout risk with continuous and categorical predictors in a Cox model: as GPA rises the dropout hazard
    drops markedly (HR = 0.569, p = 0.002), concordance 0.65 (reasonable discrimination). Cox regression gives the
    effect of several predictors on "event time" as hazard ratios (HR) in censored time data, making no assumption
    about the baseline hazard's shape. In education it is the gold standard for identifying the factors that drive
    dropout/withdrawal risk.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT)
    file: 88_aft_weibull.xlsx
  >> SCENARIO (narration):
    We model survival time assuming a Weibull distribution. For explicit time
    estimation, AFT is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_term
    - Event: graduate
    - Predictor(s): scholarship

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution: Weibull   Median survival = 6.91 terms   time: duration_term, event: graduate, predictor: scholarship

>> COMMENTARY (narration):
    We modeled survival time by assuming a parametric (Weibull) distribution: median survival ~6.9 terms. AFT
    (accelerated failure time) models, unlike Cox, choose an explicit distribution for the hazard shape and directly
    interpret how predictors "accelerate/decelerate" time. If the data fit the assumed distribution they are more
    powerful than Cox. In education they are preferred when explicit time-to-graduation estimation and extrapolation
    are needed.

====================================================================================

#89  Competing Risks
    file: 89_competing_risks.xlsx
  >> SCENARIO (narration):
    We model a setting where a student can meet several distinct ends. For mutually
    exclusive events, competing risks is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_term
    - Event: separation -> olay tipi / event type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 49   different separation types modeled as distinct event types

>> COMMENTARY (narration):
    We examined a setting where a student can meet more than one distinct end (separation types: graduation, transfer,
    dropout) in the competing-risks framework: 49 observations censored, the rest split into distinct event types.
    While standard survival treats all events alike, the competing-risks method accounts for the fact that "once one
    occurs the others no longer can" and gives a separate cumulative incidence for each event type. In education it is
    necessary to correctly model a student's mutually exclusive distinct ends.

====================================================================================

#90  Time-Dependent Cox
    file: 90_tvcox.xlsx
  >> SCENARIO (narration):
    We build a Cox model where the predictor (GPA) changes over time. For a
    time-varying covariate, time-dependent Cox is appropriate.
  >> VARIABLE SELECTION:
    - Unit (id): student_id
    - Start: start
    - Stop: end
    - Event: event
    - Predictor(s): GPA

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    GPA: HR = 1.027   p = 0.944 ns   (time-varying covariate)   id: student_id, start/stop intervals

>> COMMENTARY (narration):
    We built a Cox model where the predictor (GPA) CHANGES over time: for each student the current GPA value was used
    via start-stop intervals; the effect here is non-significant (HR = 1.027, p = 0.94). Time-dependent Cox correctly
    handles non-constant covariates (changing GPA, changing condition) -- it uses the predictor's current value at the
    event time. In education it is the right method for modeling the effect of time-varying risk factors (changing
    achievement, changing condition) on an event.

====================================================================================

#91  Survey Cox Regression
    file: 91_survey_phreg.xlsx
  >> SCENARIO (narration):
    We carry survival analysis into a complex survey design (weight+cluster). For
    design-faithful event time, survey_phreg is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_term
    - Event: dropout
    - Predictor(s): region (faktorize/factorized)
    - Weight: weight
    - Cluster: school_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.50   region_n: p = 0.863 ns   time: duration_term, event: dropout   weight: weight, cluster: school_id

>> COMMENTARY (narration):
    We carried survival/dropout analysis into a complex survey design (weight + cluster) with a Cox model: the region
    effect is non-significant (p = 0.86), concordance 0.50 (weak discrimination). survey_phreg extends the
    proportional-hazards model to a stratum/cluster/weight structure -- standard errors are corrected for the design.
    In education it is the design-faithful way to model event time (dropout, graduation) in representative
    panel/survey data.

====================================================================================

#92  Interval-Censored Survival
    file: 92_interval_censored.xlsx
  >> SCENARIO (narration):
    We model data where the event is known only within an interval. For interval
    censoring, this is appropriate.
  >> VARIABLE SELECTION:
    - Lower bound: left_censor
    - Upper bound: survival_censor

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of events = 48   Median survival = 12.0   (lower bound: left_censor, upper bound: survival_censor)

>> COMMENTARY (narration):
    We modeled data where the exact event time is unknown and only known to have occurred within an INTERVAL: median
    survival 12 units, 48 events. Interval-censored methods are for cases where the event lies "somewhere between two
    observations" (between periodic measurements); fixing the event to the interval's mid/end point biases results,
    while this method carries the uncertainty correctly. In education it is used to correctly model events occurring
    between periodic inspections (status change between two terms).

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty_cox.xlsx
  >> SCENARIO (narration):
    We model the school-specific hidden risk as frailty in clustered survival data.
    For shared hidden risk, frailty Cox is appropriate.
  >> VARIABLE SELECTION:
    - Time: duration_term
    - Event: dropout
    - Predictor(s): clinical_group (faktorize/factorized)
    - Cluster: school_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.499   cg_n: p = 0.216 ns   time: duration_term, event: dropout   cluster: school_id (frailty)

>> COMMENTARY (narration):
    In clustered survival data (students within the same school) we modeled the school-specific unobserved risk as
    "frailty" (a random effect); the clinical-group effect is non-significant (p = 0.22). Frailty Cox accounts for the
    hidden risk shared by units in the same cluster -- solving the independence assumption that standard Cox violates.
    In education it is the right choice for event data clustered within a school/class (shared hidden risk).

====================================================================================

#94  Time Series Analysis
    file: 94_ts.xlsx
  >> SCENARIO (narration):
    We examine a monthly series (trend, season, stationarity). For time-dependent
    structure, time series analysis is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: record

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.858 (not stationary)   Trend: increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly series (record): the ADF test shows it is not stationary (p = 0.86), with an increasing trend
    and seasonality. Time series analysis reveals the time-dependent structure (trend, season, autocorrelation) of
    observations; ordinary statistics are misleading because of this dependency. Stationarity is a prerequisite for
    models like ARIMA; if non-stationary, differencing is needed. In education/social research it is the starting step
    for analyzing application/enrollment/activity series.

====================================================================================

#95  STL Decomposition
    file: 95_stl.xlsx
  >> SCENARIO (narration):
    We decompose the motivation series into trend, season and residual. For seasonal
    decomposition, STL is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Season period = 7   series decomposed into trend + season + residual components

>> COMMENTARY (narration):
    We decomposed the motivation series into three components with STL: trend, the seasonal pattern (period = 7), and
    residual. STL visually and numerically separates a time series' long-term trend, recurring seasonal pattern and
    unexplained fluctuation; this answers "what is the underlying trend, and how much is seasonal?". In education/social
    research it is fundamental for de-seasonalizing series to see the underlying trend and for anomaly detection.

====================================================================================

#96  ARIMA Forecast
    file: 96_arima.xlsx
  >> SCENARIO (narration):
    We model the series with ARIMA and produce a forecast. For series forecasting,
    ARIMA is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: graduate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model selected   AIC = 2512.56   forecast + confidence interval

>> COMMENTARY (narration):
    We modeled the series (graduate) with ARIMA and produced a forecast (AIC = 2512.56 for model selection). ARIMA
    forecasts the future from the series' own past values (AR), trend (I - differencing) and past errors (MA); on a
    stationarized series it gives strong short-to-medium-term forecasts. In education/social planning it is the most
    common classical method for application/enrollment/demand forecasting; the confidence interval shows the forecast
    uncertainty.

====================================================================================

#97  Exponential Smoothing (ETS)
    file: 97_ets.xlsx
  >> SCENARIO (narration):
    We model the application series with Holt-Winters exponential smoothing. For a
    seasonal trended series, ETS is appropriate.
  >> VARIABLE SELECTION:
    - Date: date
    - Value: application

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method: Holt-Winters (seasonal additive, period = 12)   AIC = 965.83

>> COMMENTARY (narration):
    We modeled the application series with Holt-Winters exponential smoothing: it jointly estimates the level, trend
    and a 12-period seasonal component. ETS tracks the series' current level, trend and season by weighting recent
    observations more (exponentially decaying weights); for seasonal and trended series it is a practical alternative
    to ARIMA. In education/social planning it gives fast, reliable results for application/enrollment forecasting with
    regular seasonal patterns.

====================================================================================

#98  Mann-Kendall Trend
    file: 98_mann_kendall.xlsx
  >> SCENARIO (narration):
    We test a significant trend in the reading-score series without distributional
    assumptions. For robust trend detection, Mann-Kendall is appropriate.
  >> VARIABLE SELECTION:
    - Value: reading_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   significant trend detected   (trend magnitude via Sen's slope)   variable: reading_score

>> COMMENTARY (narration):
    We tested whether the reading-score series has a significant trend -- without assuming a distribution -- with
    Mann-Kendall: p < .001, a significant trend; Sen's slope gives the robust (outlier-resistant) magnitude of the
    trend. Mann-Kendall is a nonparametric trend test; it requires no normality and is resistant to outliers, hence
    common in time-series work. In education it is used to robustly detect long-term indicator trends (rising/falling
    achievement trend).

====================================================================================

#99  Anomaly Detection
    file: 99_anomali.xlsx
  >> SCENARIO (narration):
    We detect unusual observations in multivariate data. For composite outlier
    detection, anomaly detection is appropriate.
  >> VARIABLE SELECTION:
    - Variables: GPA
    - Variables: study
    - Variables: absent
    - Variables: score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 18   variables: GPA, study, absent, score

>> COMMENTARY (narration):
    We automatically detected unusual observations in multivariate data: 18 students were flagged as anomalies. Anomaly
    detection finds observations that deviate markedly from normal (erroneous record, exceptional case) by evaluating
    several variables together; it catches "composite" outliers that univariate thresholds miss. In education it is
    used for data-quality control, erroneous-record detection and the early detection of exceptional student profiles.

====================================================================================

#100  Variance Components
    file: 100_varcomp_3seviye_h2.xlsx
  >> SCENARIO (narration):
    We decompose measurement variability into nested levels (upper/lower unit). For
    hierarchical variability, variance components is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): upper_unit
    - Factor (categorical): lower_unit (nested)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Random factors: upper_unit + lower_unit (nested within upper)   % contribution of each component + residual

>> COMMENTARY (narration):
    We decomposed a measurement's variability into nested levels -- upper unit and lower unit within upper: the percent
    contribution of each level and the residual were reported. Variance components analysis answers "how much of the
    variability is between upper units, how much between lower units, how much within unit?". In education it is used
    in hierarchical structures (school>class>student) to see where uncertainty concentrates and in sampling design.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old.xlsx
  >> SCENARIO (narration):
    We examine two groups' measurement difference with a Bayesian t-test, expressing
    evidence as a Bayes factor. For an intuitive evidence ratio, this is
    appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Grouping (categorical): group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 3.36e+46   Cohen's d = 3.939   (overwhelming evidence for H1)   outcome: value, group: group

>> COMMENTARY (narration):
    We examined the measurement difference between two groups (new/old method) with a Bayesian t-test: the Bayes factor
    is overwhelming (BF10 ~ 3.4e46), the effect very large (d = 3.94) -- the data support the "difference" hypothesis
    over "no difference" by astronomical odds. Unlike a p-value, the Bayes factor gives the RELATIVE evidence strength
    of two hypotheses and can distinguish "no evidence" from "no difference". In education it is preferred when one
    wants to express the evidential strength of a decision as an intuitive ratio.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BF10.xlsx
  >> SCENARIO (narration):
    We evaluate the relationship between two variables in a Bayesian framework. For
    evidential strength, Bayesian correlation is appropriate.
  >> VARIABLE SELECTION:
    - 1st measure / group: X_variable
    - 2nd measure / group: Y_variable

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.780 (very strong)   BF10 = 4.50e+14   (overwhelming evidence for a relationship)

>> COMMENTARY (narration):
    We evaluated the relationship between two variables in a Bayesian framework: r = 0.78 and BF10 ~ 4.5e14, i.e. very
    strong evidence for a relationship. Bayesian correlation, instead of a classical p-value, presents the evidential
    strength of the relationship as a Bayes factor and the coefficient's posterior distribution. In education it is
    valuable for reporting not just whether the relationship between two indicators is "significant" but how strongly
    it is "evidenced".

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_2yonlu.xlsx
  >> SCENARIO (narration):
    We examine a measurement's difference across two factors with Bayesian ANOVA.
    For the evidential strength of factor effects, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): factor1
    - 2nd Factor: factor2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor1: BF10 = 76099 (very strong evidence)   outcome: value, factors: factor1, factor2

>> COMMENTARY (narration):
    We examined a measurement's difference across two factors with Bayesian ANOVA: for factor1, BF10 ~ 76,000, very
    strong evidence. Bayesian ANOVA compares the effects of factors and their interactions via Bayes factors; it ranks
    probabilistically which model (which effects) best explains the data. Unlike classical ANOVA's "reject/don't
    reject" decision, it quantifies the relative support among models. In education it is used to compare the
    evidential strength of factor effects.

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_LMM.xlsx
  >> SCENARIO (narration):
    We analyze group-nested data with a Bayesian hierarchical model. For stable
    estimates in small groups, this is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_response
    - Cluster: group_id
    - Predictor(s): X_covariate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    sigma^2_u (group) = 2.90 (8.6%)   sigma^2_eps (residual) = 30.79 (91.4%)   outcome: Y_response, group: group_id

>> COMMENTARY (narration):
    We analyzed group-nested data with a Bayesian hierarchical (multilevel) model: 8.6% of variability is
    between-group, 91.4% within-group (ICC ~ 0.09). The Bayesian hierarchical model is the Bayesian version of LMM -- it
    estimates group effects with posterior distributions and balances small groups via "partial pooling". In education
    it is powerful for producing stable estimates even in small groups within multilevel (school/class/student) data.

====================================================================================

#105  Spatial SAR
    file: 105_spatial_sar_spatial.xlsx
  >> SCENARIO (narration):
    When modeling a measurement we handle spatial spillover with SAR. For
    neighborhood effects, SAR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho = 0.336   z = 4.94   p < .001 ***   Pseudo R^2 = 0.87   (N = 100, k-NN W, k = 5)
    outcome: Y_value, predictors: X1, X2

>> COMMENTARY (narration):
    When modeling a measurement (Y_value), we handled spatial spillover (the effect of neighboring units) with SAR: the
    spatial lag parameter is significant and positive (rho = 0.34, p < .001) -- a unit's value is related to its
    neighbors' value, a "cluster/spillover" pattern. SAR incorporates spatial dependency into the model; if ignored,
    standard errors are biased. In education/social geography it is the right method for modeling the geographic spread
    of school/region indicators (neighborhood effect).

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_residual.xlsx
  >> SCENARIO (narration):
    We model spatial dependency in the error term. For unmeasured geographic
    factors, SEM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda = 0.590   z = 6.38   p < .001 ***   Pseudo R^2 = 0.84   (N = 100, k-NN W, k = 5)

>> COMMENTARY (narration):
    This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
    0.59, p < .001) -- the effect of geographic variables omitted from the model makes neighboring errors correlated.
    Unlike SAR, SEM attributes the spread to the error rather than the outcome. In education/social geography, when the
    source of spatial autocorrelation is unmeasured geographic factors (socioeconomic environment, access), the correct
    specification is SEM; it is chosen by comparison with SAR.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local.xlsx
  >> SCENARIO (narration):
    Assuming the relationship is not constant in space, we estimate separate
    coefficients per location. For spatial heterogeneity, GWR is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictor(s): X1
    - Predictor(s): X2
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R^2 = 0.887   local coefficients vary by location   outcome: Y_value, predictors: X1, X2

>> COMMENTARY (narration):
    Assuming the relationship is NOT constant across space, we estimated SEPARATE coefficients for each location (R^2 =
    0.89). Unlike "global" regression, GWR fits a separate model at each point with local overlap/weights; this answers
    "how does this variable's effect vary by region?" and maps spatial heterogeneity. In education/social geography it
    is powerful where the education-socioeconomic relationship differs by region.

====================================================================================

#108  Nested Mixed Model
    file: 108_nested_lmm_R_P_F.xlsx
  >> SCENARIO (narration):
    We model a value in a nested design (upper>lower>block). For nested hierarchies,
    nested LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): lower_group_no
    - Factor (categorical): upper_group
    - Factor (categorical): block

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper_group (P) effect: F = 27.44, p < .001 ***   lower_group_no (R) effect: F = 6.11, p < .001 ***
    nested variance components (block within P)

>> COMMENTARY (narration):
    We modeled a value in a nested design (upper group > lower group > block): both the upper level (F = 27.44, p <
    .001) and the lower/replication level (F = 6.11, p < .001) make significant contributions. Nested LMM correctly
    handles hierarchies where sub-units are nested within super-units (each block belongs to only one group); by
    partitioning variance into levels it shows each layer's share. In education it is used to correctly separate
    effects in school>class>student hierarchies.

====================================================================================

#109  Crossed Mixed Model
    file: 109_crossed_lmm_A_B.xlsx
  >> SCENARIO (narration):
    We model a design where two random factors are crossed. For two independent
    classification axes, crossed LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Factor (categorical): factor_a
    - 2nd Factor: factor_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    A x B interaction: F = 14.67, p < .001 ***   main effects ns (A: p=0.767, B: p=0.480)   outcome: value

>> COMMENTARY (narration):
    We modeled a design where two random factors are crossed rather than nested (each level of A pairs with each level
    of B): while the main effects are non-significant (A: p=0.77, B: p=0.48), the A x B interaction is very strong
    (F = 14.67, p < .001) -- so the effect depends on the COMBINATION of factors. Crossed LMM, unlike nested, handles
    two independent grouping axes (e.g. rater x item, each rater on each item) at once. In education/psychometrics it
    is the right choice for jointly analyzing the effects and interaction of two independent classification axes.

====================================================================================

#110  Kernel Density (KDE) Map
    file: 110_KDE_traffic_density.xlsx
  >> SCENARIO (narration):
    We produce a continuous density surface from unit locations. For a density map,
    KDE is appropriate.
  >> VARIABLE SELECTION:
    - Value: monthly_income_TL
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 95 units   income-weighted spatial density surface (continuous heat map)

>> COMMENTARY (narration):
    Using unit locations (n = 95) we produced a continuous density surface (a KDE heat map): the point distribution was
    turned into a smooth "density" surface. KDE answers "where is it dense?" from scattered point data as a continuous
    map; it shows the trend rather than individual points. In education/social geography it is the core spatial tool
    for visually mapping student/school/socioeconomic concentrations and identifying empty/dense zones.

====================================================================================

#111  Hexbin Density Map
    file: 111_Hexbin_measurement_noktalari.xlsx
  >> SCENARIO (narration):
    We aggregate unit distribution into hexagonal cells to show density. For the
    over-plotting problem, hexbin is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 150 units   spatial density via aggregation into hexagonal cells

>> COMMENTARY (narration):
    We colored unit distribution (n = 150) by counting within hexagonal cells. Hexbin aggregates many overlapping
    points (the over-plotting problem) into a regular hexagonal grid to show density clearly; compared with squares it
    carries less directional bias. In education/social geography it is used to turn dense point clouds (student/school
    locations) into a readable density map and to compare spatial concentrations.

====================================================================================

#112  Moran's I
    file: 112_Morans_I_structure_quality_autocorrelation.xlsx
  >> SCENARIO (narration):
    We test whether income is distributed randomly or in clusters across space. For
    spatial autocorrelation, Moran's I is appropriate.
  >> VARIABLE SELECTION:
    - Value: monthly_income_TL
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.467   z = 8.31   p < .001 ***   (k-NN neighbors, k = 6)   variable: monthly_income_TL

>> COMMENTARY (narration):
    We tested whether income is distributed randomly or in clusters across space with Moran's I: I = 0.47, z = 8.31,
    p < .001 -- strong positive spatial autocorrelation, i.e. similar income levels cluster geographically (the wealthy
    with the wealthy, the lower-income together). Moran's I quantifies "Tobler's first law" (near things are similar).
    In education/social geography it is the first test for detecting the geographic clustering of socioeconomic
    variables and for deciding whether a spatial model is needed.

====================================================================================

#113  Getis-Ord Gi*
    file: 113_Getis_Ord_traffic_hotspot.xlsx
  >> SCENARIO (narration):
    We map locally where income clusters high/low. For hot/cold spots, Getis-Ord is
    appropriate.
  >> VARIABLE SELECTION:
    - Value: monthly_income_TL
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Hot spots = 27   Cold spots = 27   n = 68   (k-NN k = 8, binary weights)   variable: monthly_income_TL

>> COMMENTARY (narration):
    We mapped locally WHERE income clusters high/low with Getis-Ord Gi*: 27 significant hot spots (high-income
    clusters) and 27 cold spots (low-income clusters). While Moran's I states the overall clustering, Gi* shows its
    location -- for each point it tests "is the surrounding area high or low?". In education/social policy it is used to
    pinpoint high/low-socioeconomic zones (hot/cold spots) for targeted intervention/resource allocation.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_structure_kumeleri.xlsx
  >> SCENARIO (narration):
    We cluster unit locations with density-based DBSCAN. For spatial clusters and
    outlier locations, DBSCAN is appropriate.
  >> VARIABLE SELECTION:
    - Latitude: lat
    - Longitude: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Clusters = 4   Noise (outliers) = 11   n = 115   (location: lat, lon)

>> COMMENTARY (narration):
    We clustered unit locations (n = 115) with density-based DBSCAN: 4 natural geographic clusters and 11 "noise"
    (scattered, cluster-less) points. DBSCAN does not require the cluster count in advance, finds clusters of any shape,
    and separates sparse points as outliers -- ideal for spatial clustering. In education/social geography it is used
    to detect neighborhood/region-level natural concentrations (student/school clusters) and to isolate isolated
    locations; it completes our descriptive spatial analysis series.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_attitude_score.xlsx
  >> SCENARIO (narration):
    We follow 40 participants measured at three phase levels (before, after, followup); each belongs to one of two group groups (intervention / control). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on attitude score.
  >> VARIABLE SELECTION:
    - Dependent variable: attitude_score
    - Subject ID: participant_id
    - Between-subjects factor: group
    - Within-subjects factor: phase

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (group): F(1,38) = 2.76  p = 0.105  np2 = 0.068
    Within (phase):  F(2,76) = 89.05  p < .001  np2 = 0.701
    Interaction:         F(2,76) = 10.49  p < .001  np2 = 0.216
    Mauchly W = 0.999  p = 0.987   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a group x phase mixed design we analyzed attitude score for 40 participants (120 observations). The interaction is significant (F(2,76) = 10.49, p < .001, np2 = 0.216) *** -- the two groups' change across phase differs in magnitude. The between-subjects main effect (intervention vs control) is F = 2.76, p = 0.105; the within-subjects main effect (before/after/followup) is F = 89.05, p < .001. Mauchly's test p = 0.987, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each phase level. In social sciences, the mixed design is the standard analysis for intervention-vs-control attitude change over phases.

====================================================================================
