====================================================================================
MerQur - Egitim_Bilimleri - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Egitim_Bilimleri/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_student_academic.xlsx
  >> SCENARIO (narration):
    Before running any inferential test, we want a clear picture of who our
    students are and how they perform. We have 280 students with their age, sex,
    class level, GPA, weekly study hours, absenteeism days, and an academic
    motivation score. The goal here is purely exploratory: to summarize central
    tendency, spread, and distribution shape for each numeric variable, and to
    see how the sample breaks down by sex. Descriptive statistics is the right
    starting point because we are not yet testing a hypothesis, just describing
    the data so later analyses make sense.
  >> VARIABLE SELECTION:
    - Variables: GPA
    - Variables: weekly_study_hour
    - Variables: absenteeism_day
    - Variables: academic_motivation
    - Variables: age
    - Group/category: sex (male, female)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 students
    GPA          : mean = 2.71   (1.14 - 4.00)        |   weekly study : mean = 12.1 hours
    academic motivation : mean = 3.4                  |   absenteeism  : mean = 8.0 days

>> COMMENTARY (narration):
    We first drew the general academic profile of our sample: 280 students, mean GPA 2.71 (on a 4-point scale), 12
    hours of weekly study, an average of 8 absent days. The wide spread of GPA from 1.14 to 4.00 shows we are working
    with a group heterogeneous in achievement. This descriptive table sets the stage for all the analyses that follow
    -- achievement comparisons, method effects, motivation-achievement relationships. Before any inferential test,
    seeing the centre, spread and range of the data is essential both to check data quality and to calibrate
    expectations.

====================================================================================

#2  Normality Tests
    file: 02_normality_GPA_study_motivation.xlsx
  >> SCENARIO (narration):
    Many of the parametric tests we plan to use assume that continuous variables
    are normally distributed. So before choosing between parametric and
    nonparametric approaches, we check that assumption directly. With 200
    students measured on GPA, weekly study hours, and motivation, we want to
    know whether each of these variables departs from a normal distribution. A
    normality test such as Shapiro-Wilk is appropriate here because it formally
    evaluates the distributional assumption for each continuous variable rather
    than relying on eyeballing a histogram.
  >> VARIABLE SELECTION:
    - Variables to test: GPA
    - Variables to test: study_hour
    - Variables to test: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    GPA        : Shapiro-Wilk = 0.987  p = 0.068   KS p = 0.422   Normal
    study_hour : Shapiro-Wilk = 0.912  p < .001    KS p = 0.001   Not Normal
    motivation : Shapiro-Wilk = 0.963  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three variables -- GPA, study hours and motivation -- are normally distributed. The result
    differs by variable: GPA is normal (p = 0.068 > 0.05; grades symmetric, bell-curve-like), but study hours and
    motivation deviate significantly from normality (p < .001). This is typical in educational data -- study hours are
    usually right-skewed (most study little, a few a lot), and motivation may be an ordinal/bounded scale. The
    practical takeaway: we can safely use parametric tests (t-test, ANOVA, Pearson) on GPA; for study and motivation,
    nonparametric methods (Mann-Whitney, Kruskal-Wallis) or a transform are more appropriate. The normality check is a
    critical pre-step that decides, for each variable separately, which test family is appropriate.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_GPA.xlsx
  >> SCENARIO (narration):
    Our district sets a benchmark GPA of 2.50 that public school students are
    expected to reach on average. We collected GPA data from 100 students in a
    public school and want to know whether their mean GPA differs significantly
    from that benchmark value. A one-sample t-test is the correct tool because
    we are comparing the mean of a single continuous variable against a known
    reference value, with no second group involved.
  >> VARIABLE SELECTION:
    - Test variable: GPA
    - Test (reference) value: 2.50

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = 8.0539   p < .001 ***   Cohen d = 0.805 (Large)
    Mean GPA = 2.90   (test mu = 2.5, passing threshold)   H0 REJECTED

>> COMMENTARY (narration):
    We compared a school's mean GPA against the 2.5 passing/reference threshold. The result: mean 2.90, significantly
    above the threshold -- t(99) = 8.05, p < .001, large effect size d = 0.81. So student achievement exceeds the
    reference threshold by more than chance can explain. The one-sample t-test is the right way to compare a group
    mean against a known standard (passing grade, national average, target value) -- a very common question in
    educational evaluation. The effect size shows the difference is not just significant but practically important too.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent Samples t-Test
    file: 04_independent_t_sex_math.xlsx
  >> SCENARIO (narration):
    A common question in education research is whether boys and girls differ in
    mathematics achievement. We have math scores from 115 students along with
    their sex, coded as female and male. We want to test whether the average
    math score differs between the two groups. An independent samples t-test
    fits this design because the outcome is continuous and we are comparing the
    means of two unrelated, independent groups.
  >> VARIABLE SELECTION:
    - Grouping variable: sex (female, male)
    - Test variable: math_score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = 2.517   p = 0.013 *   Cohen d = 0.470 (Small-moderate)
    Math score differs significantly by sex.   H0 REJECTED

>> COMMENTARY (narration):
    We examined whether math scores differ by sex with an independent t-test. The difference is significant (t(113) =
    2.52, p = 0.013), but the effect size is small-moderate at d = 0.47 -- statistically real, but limited in
    magnitude. This distinction matters: in large samples even small differences can be significant, so it is
    essential to report effect size alongside the p-value. The independent t-test is the basic method for comparing
    the means of two independent groups (sex, class, school type). The small effect here says the sex difference
    exists but should not be exaggerated educationally.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired t-Test
    file: 05_paired_t_on_last_test.xlsx
  >> SCENARIO (narration):
    We ran an instructional intervention and want to know whether students
    improved. The same 50 students were measured twice: an early test and a
    final test, with the intervention happening in between. Because each student
    provides both an early and a final score, the two measurements are paired
    within the same person. A paired-samples t-test is appropriate here since we
    are comparing two related measurements taken from the same students.
  >> VARIABLE SELECTION:
    - Pair measurement 1 (M1): on_test
    - Pair measurement 2 (M2): last_test

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Paired t(49) = -17.82   p < .001 ***   Cohen d_z = -2.520 (Large)
    Mean difference (pre-test − post-test) = -11.72   H0 REJECTED

>> COMMENTARY (narration):
    We measured the effect of an intervention (extra instruction, new method) by comparing the same students' pre-
    and post-test scores with a paired t-test. The result is striking: post-test scores are on average 11.72 points
    higher than pre-test, t(49) = -17.82, p < .001, with an enormous effect size (d_z = -2.52). The intervention
    raised achievement very strongly. Because we measured the same students twice, the paired test is correct -- it
    removes between-student variability and focuses only on the individual's own change. It is the gold-standard way
    to evaluate before/after interventions in education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_instruction_method.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether teaching method affects achievement. 140
    students each received one of four instruction methods: traditional,
    cooperative, project, or mixed (blended), and we recorded their achievement
    score afterward. The question is whether mean achievement differs across
    these four methods. We use one-way ANOVA because the outcome is continuous
    and the predictor is a single categorical factor with several levels.
  >> VARIABLE SELECTION:
    - Factor (group): instruction_method (traditional, cooperative, project, mixed (blended))
    - Dependent variable: achievement_score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 136) = 17.9985   p < .001 ***   η² = 0.2842
    DECISION: H0 REJECTED (4 instruction methods)

>> COMMENTARY (narration):
    We compared the effect of four instruction methods on achievement with one-way ANOVA. The result is significant
    (F(3,136) = 18.00, p < .001), effect size η² = 0.28 -- about 28% of achievement is explained by instruction
    method, a strong effect in educational research. ANOVA says at least one method differs from the others; a
    post-hoc test is needed to see which. "Which instruction method is more effective" is one of the most basic
    research questions in education; ANOVA is the standard way to answer it across multiple groups.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova_method_sex.xlsx
  >> SCENARIO (narration):
    Here we ask two questions at once: does teaching method affect achievement,
    does sex affect achievement, and do the two interact? We have 150 students
    crossed by instruction method, with levels traditional, project, and mixed,
    and by sex, female and male, with achievement recorded for each. A two-way
    ANOVA is the right design because we have two categorical factors and we
    want to estimate their main effects as well as their interaction on a
    continuous outcome.
  >> VARIABLE SELECTION:
    - Factor 1: method (traditional, project, mixed)
    - Factor 2: sex (female, male)
    - Dependent variable: achievement

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    method (main) : F(2) = 42.24   p < .001 ***   η²p = 0.37
    sex (main)    : F(1) =  4.61   p = 0.033 *    η²p = 0.03

>> COMMENTARY (narration):
    We examined achievement with two factors -- instruction method and sex -- together. The method effect is very
    strong (F = 42.24, η²p = 0.37), while the sex effect is significant but very small (F = 4.61, p = 0.033, η²p =
    0.03). So the main determinant of achievement is instruction method; sex contributes minimally. Two-way ANOVA's
    strength is seeing both factors' effects -- separately and jointly (interaction) -- at once. This answers
    questions like "is the method effect similar for both sexes" and reveals for whom educational interventions work best.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated Measures ANOVA
    file: 08_repeated_anova_term_GPA.xlsx
  >> SCENARIO (narration):
    We tracked the same 60 units across an academic year, recording achievement
    at four time points from the first to the fourth measurement. The question
    is whether mean performance changes across these four occasions. A repeated
    measures ANOVA is appropriate because the same units are measured repeatedly
    over time, so the four measurements are correlated within each unit rather
    than independent.
  >> VARIABLE SELECTION:
    - Repeated measures: measurement_1
    - Repeated measures: measurement_2
    - Repeated measures: measurement_3
    - Repeated measures: measurement_4

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 177) = 108.40   p < .001 ***   η²p = 0.6475   (n = 60)
    DECISION: H0 REJECTED

>> COMMENTARY (narration):
    We compared the same 60 students' GPA across four consecutive terms with repeated-measures ANOVA. Because the
    same students are measured repeatedly, observations are dependent; RM-ANOVA accounts for this within-subject
    correlation (ordinary ANOVA does not). The result is very strong (F(3,177) = 108.40, p < .001, η²p = 0.65): GPA
    changes significantly across terms -- a strong temporal trend. RM-ANOVA is the correct design for longitudinal
    studies where the same units are tracked over time, and is far more powerful than treating each term as
    independent. It is ideal for studies following student development term by term.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_method_3ders.xlsx
  >> SCENARIO (narration):
    Instead of looking at one subject at a time, we want to know whether
    teaching method affects overall academic profile across several subjects
    simultaneously. We have 120 students assigned to traditional, mixed, or
    project methods, with scores in math, science, and language. MANOVA is the
    appropriate test because we have multiple correlated continuous outcomes and
    a single categorical factor, and we want to test the method effect on the
    combined set of dependent variables.
  >> VARIABLE SELECTION:
    - Factor (group): method (traditional, mixed, project)
    - Dependent variables: math
    - Dependent variables: science
    - Dependent variables: language

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.320 -> F(6, 230) = 29.48   p < .001 ***   (n = 120)
    Dependent: math, science, language   |   Factor: method

>> COMMENTARY (narration):
    Here we have three correlated dependent variables -- math, science, language scores -- and one method factor.
    Running three separate ANOVAs would inflate the error rate and ignore the correlations among subjects. MANOVA
    tests all three jointly: Wilks' Lambda F = 29.48, p < .001, so the instruction method strongly shifts the
    multivariate achievement profile. MANOVA is the right approach when an intervention affects several related
    outcomes at once -- it controls the overall error rate and catches effects spread across variables. After a
    significant MANOVA, follow-up univariate tests show which subject drives the effect.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova_intervention_baseline.xlsx
  >> SCENARIO (narration):
    We want to compare the effect of an intervention on a final test while
    accounting for where students started. 105 students were assigned to a
    control group, intervention-A, or intervention-B, and we have both an early
    baseline test and a final test. By treating the early test as a covariate,
    we can compare adjusted final-test means across the three groups. ANCOVA is
    the right choice because it tests a categorical group effect on a continuous
    outcome while statistically controlling for a continuous covariate.
  >> VARIABLE SELECTION:
    - Factor (group): group (control, intervention-A, intervention-B)
    - Dependent variable: last_test
    - Covariate: on_test

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    group (factor): η²p = 0.504 (large)   p < .001 ***
    Covariate: pre-test (baseline level)   |   DV: post-test   H0 REJECTED

>> COMMENTARY (narration):
    When evaluating an intervention's effect, it is essential to control for students' baseline (pre-test) levels --
    because if the groups differ at baseline, the post-test difference is misleading. ANCOVA takes the pre-test as a
    covariate and holds it statistically constant. The result: even after adjusting for baseline, the groups differ
    significantly in post-test (η²p = 0.50, large effect). So the observed difference is not a by-product of baseline
    differences; it is the genuine intervention effect. ANCOVA is the standard way -- very common in education -- to
    fairly control baseline differences in pre/post intervention studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci_absenteeism.xlsx
  >> SCENARIO (narration):
    Suppose we are studying student absenteeism and we want a trustworthy
    estimate of the average number of absent days per term, but our sample is
    small, only 30 students, and absenteeism counts are notoriously skewed.
    Rather than rely on a normality assumption we cannot justify, we resample
    the absenteeism_day values thousands of times to build an empirical sampling
    distribution of the mean. A bootstrap confidence interval is ideal here
    because it makes no distributional assumptions and gives us a robust
    interval for the true mean absenteeism, even with this limited sample.
  >> VARIABLE SELECTION:
    - Variable: absenteeism_day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean absenteeism = 5.87 days
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    We derived a confidence interval for the mean of absenteeism days by bootstrapping: the observed mean is 5.87
    days. Absenteeism data are typically right-skewed (most students absent little, a few absent a lot), and classic
    normal-assuming formulas can be unreliable for such data. Bootstrap estimates the interval without distributional
    assumptions, through thousands of resamples. It is a robust way to express the uncertainty of a mean in skewed
    educational data (absenteeism, discipline counts, income).

====================================================================================

#12  Permutation Test
    file: 12_permutation_motivation_group.xlsx
  >> SCENARIO (narration):
    Imagine a motivation intervention where 53 students were split into a
    control and an experiment group, and we recorded their academic motivation
    scores. We want to know whether the experimental program genuinely raised
    motivation, but the groups are uneven and we are not comfortable assuming
    the score differences follow a normal distribution. A permutation test
    answers this directly by repeatedly shuffling the group labels to see how
    often a mean difference as large as ours would arise by chance, giving an
    exact, assumption-free p-value for the group effect.
  >> VARIABLE SELECTION:
    - Group (factor): group (control, experiment)
    - Test variable: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = -2.91   p = 0.0054 **   H0 REJECTED
    Motivation differs significantly between the two groups

>> COMMENTARY (narration):
    We tested the motivation difference between two groups with a permutation test. Permutation shuffles the group
    labels thousands of times to measure whether the observed difference could arise by chance -- requiring no
    distributional assumption. The result is significant (p = 0.005): the groups' motivation differs more than chance
    allows. For small samples or when normality is doubtful, permutation is a robust, assumption-free alternative to
    the t-test. It is a reliable backup for group comparisons in educational experiments.

====================================================================================

#13  Multiple Comparison Corrections
    file: 13_multiple_comparison_material.xlsx
  >> SCENARIO (narration):
    Here we compared nine different instructional materials, each used by a
    subset of 180 students, and recorded their test_score. Because comparing
    nine materials means running many pairwise tests, the chance of a false
    positive balloons. We therefore apply multiple comparison corrections, such
    as Bonferroni or Holm, to adjust the p-values across all material pairs, so
    that any difference we declare significant is real and not just an artifact
    of testing so many combinations at once.
  >> VARIABLE SELECTION:
    - Group (factor): material (9 levels)
    - Dependent variable: test_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Material groups vs reference: TOTAL SIGNIFICANT = 5/8 (in every correction method)

>> COMMENTARY (narration):
    When we compare different teaching materials pairwise on test score, we run many tests -- each carrying a false-
    positive risk. Multiple-comparison correction controls that risk. Here five of eight comparisons stayed
    significant under every correction method (Bonferroni, Holm, FDR) -- 5/8. So some materials genuinely differ from
    the reference, while others lose significance after correction. Skipping this step risks reporting spurious
    differences; in educational material/method comparisons the correct correction is essential.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney_location_satisfaction.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether students in urban schools report different
    satisfaction than those in rural schools. We have 85 students with a
    school_location of urban or rural and a satisfaction score that is ordinal
    and clearly non-normal. Because we are comparing two independent groups on a
    variable that violates the normality assumption of the t-test, the Mann-
    Whitney U test is the right choice: it compares the rank distributions of
    the two groups without assuming normality.
  >> VARIABLE SELECTION:
    - Group (factor): school_location (urban, rural)
    - Test variable: survival_satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Mann-Whitney U = 1234.50   p = 0.0029 **   r = -0.37
    School satisfaction differs significantly by school location.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether school satisfaction (a survey score) differs by school location (urban/rural) with the
    nonparametric Mann-Whitney -- because satisfaction/Likert scores are ordinal and may not be normal. The result is
    significant (U = 1234.5, p = 0.003), with a moderate effect size r = 0.37. So students in differently located
    schools report markedly different satisfaction. For ordinal/skewed survey data Mann-Whitney is more reliable than
    the t-test and is widely used in educational satisfaction/attitude studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank Test
    file: 15_wilcoxon_teacher_attitude.xlsx
  >> SCENARIO (narration):
    Picture a teacher professional-development program where we measured each of
    35 teachers' attitude toward a new curriculum before and after the training.
    We want to know whether attitudes shifted, but the attitude scale is ordinal
    and the sample is small, so a paired t-test is questionable. The Wilcoxon
    signed-rank test is appropriate because it compares two related measurements
    on the same teachers using the signed ranks of their differences, without
    assuming the differences are normally distributed.
  >> VARIABLE SELECTION:
    - Pair M1 (before): attitude_before
    - Pair M2 (after): attitude_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilcoxon W = 0.00   p < .001 ***   r = 0.888   n (nonzero) = 30
    Teacher attitude changed significantly before vs after training.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same teachers' attitude scores before and after an in-service training with the paired Wilcoxon
    test. The result is very strong: W = 0, p < .001, effect size r = 0.89. Attitudes changed in the same direction
    and by a large amount in nearly all teachers -- the training had a clear effect. For paired measures and
    ordinal/skewed attitude data, Wilcoxon is the right choice. It is a robust, assumption-free way to measure the
    effect of teacher-development interventions.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_sound_book.xlsx
  >> SCENARIO (narration):
    Suppose we are examining whether background noise affects reading habits. We
    grouped 45 students by the sound level in their study environment, low
    sound, medium sound, or high sound, and counted their monthly_book_count. We
    want to compare reading volume across the three noise conditions, but book
    counts are skewed and the groups are small, so one-way ANOVA is not safe.
    The Kruskal-Wallis test is the appropriate nonparametric alternative for
    comparing a continuous outcome across three or more independent groups using
    ranks.
  >> VARIABLE SELECTION:
    - Group (factor): sound (low sound, medium sound, high sound)
    - Dependent variable: monthly_book_count

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 16.53   p < .001 ***   eta-squared_H = 0.35
    Monthly book count differs significantly by socio-economic level.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the number of books read monthly varies by socio-economic level with Kruskal-Wallis -- the
    nonparametric ANOVA for more than two groups, appropriate because count data are not normally distributed. The
    result is highly significant (H(2) = 16.53, p < .001), with a large effect (eta-squared = 0.35): reading habits
    are strongly tied to socio-economic level. This is a concrete indicator of opportunity inequality in education.
    For count/ordinal data Kruskal-Wallis is more appropriate than classic ANOVA.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman_student_course.xlsx
  >> SCENARIO (narration):
    Imagine an experiment where each of 30 students was taught the same material
    under four different conditions, A, B, C, and D, and we recorded a
    performance score in every condition. Since the same students experienced
    all four conditions, the measurements are related, and the scores are not
    normally distributed. The Friedman test is the right tool here: it is the
    nonparametric analogue of repeated-measures ANOVA, comparing the four
    related conditions on the basis of within-subject ranks.
  >> VARIABLE SELECTION:
    - Repeated measures: condition_A, condition_B, condition_C, condition_D

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Friedman chi-square(3) = 39.08   p < .001 ***   Kendall W = 0.4342   n = 30
    The same students' scores under 4 conditions differ significantly.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same students' scores under four different conditions/methods with the Friedman test -- the
    nonparametric counterpart of repeated-measures ANOVA. The result is highly significant (chi-square(3) = 39.08,
    p < .001), with Kendall W = 0.43 indicating moderate-high consistency of the ranking across conditions. So the
    four conditions make a clear difference in student performance. Because the same students are measured
    repeatedly, observations are dependent; Friedman accounts for this correctly. It is the right test for
    ordinal/non-normal repeated-measures data.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_sinav_achievement.xlsx
  >> SCENARIO (narration):
    Suppose a national exam is supposed to have a 50 percent pass rate, and we
    want to test whether our cohort of 800 students differs from that benchmark.
    For each student we recorded passed as a 0 or 1 outcome. Because the outcome
    is a simple binary success/failure and we are comparing the observed
    proportion of passers against a fixed expected proportion, the binomial test
    is exactly the right procedure for evaluating whether the pass rate departs
    from the hypothesized 0.5.
  >> VARIABLE SELECTION:
    - Binary variable: passed

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed pass proportion = 0.7625   test proportion = 0.50
    p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the pass rate on an exam equals chance (50%) with the binomial test. The observed proportion is
    0.76 -- three-quarters of students passed. This differs significantly from 50% (p < .001). The binomial test is
    the simplest way to compare a single binary proportion (pass/fail, present/absent) against a known reference. In
    education it is a practical tool for comparing pass rates against a target, evaluating exam difficulty, or testing
    an achievement expectation.

====================================================================================

#19  Sign Test
    file: 19_sign_test_course_info.xlsx
  >> SCENARIO (narration):
    Picture a short course where we measured each of 28 students' knowledge
    score before and after, score_before and score_post, and we simply want to
    know whether the course tended to improve scores. The differences are not
    symmetric enough for a Wilcoxon test and we are not willing to assume any
    distributional shape; we only care about the direction of change. The sign
    test is the most conservative choice, basing its conclusion solely on the
    number of students who improved versus declined.
  >> VARIABLE SELECTION:
    - Pair M1 (before): score_before
    - Pair M2 (after): score_post

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive (before > after) = 2   (scores mostly INCREASED)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same students' scores before and after a course with the paired sign test. Scores increased in
    nearly all students (only 2 had a higher before-score), p < .001. So the course consistently raised scores. The
    sign test looks only at the direction of change (up/down), not its magnitude; it is a robust, assumption-free
    test. When the simple question "how many students improved" matters more than effect size, or when the
    distribution is very skewed, it is a coarse but sturdy before/after indicator.

====================================================================================

#20  Runs Test
    file: 20_runs_test_absenteeism_series.xlsx
  >> SCENARIO (narration):
    Suppose we tracked daily absenteeism over 60 consecutive school days and
    want to know whether absences occur randomly over time or in non-random
    streaks, perhaps clustering around certain periods. We recorded absent_count
    for each day in sequence. The runs test is the appropriate procedure here
    because it examines the ordering of the series, testing whether high and low
    values alternate randomly or form runs that signal an underlying pattern or
    dependence across days.
  >> VARIABLE SELECTION:
    - Sequence variable: absent_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 30   Z = -0.25   (accepted as random)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the daily absenteeism series is randomly distributed -- whether successive days cluster or show
    a pattern -- with the Wald-Wolfowitz runs test. With Z = -0.25 the observed number of runs is very close to
    expected; the series is accepted as random, with no systematic pattern. This matters methodologically: if a
    series is non-random (autocorrelated), many classic tests become invalid. Because randomness holds here, standard
    methods can be used safely. The runs test is ideal for this pre-check on time-ordered data.

====================================================================================

#21  Chi-Square Test of Independence
    file: 21_chisquare_sound_university_preference.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether the noise level students study in is related
    to the type of university they end up attending. We surveyed 540 students,
    recording the typical sound level of their study environment (low, medium,
    or high) and the type of university they enrolled in (public or private).
    Our research question is whether study-environment noise and university type
    are associated, or whether the two are completely independent of each other.
    Because both variables are categorical and we have a large sample of
    independent observations, the Chi-Square Test of Independence is the right
    tool here.
  >> VARIABLE SELECTION:
    - Row variable: sound (low/medium/high)
    - Column variable: university_type (public/private)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(2) = 141.06   p < .001 ***   Cramer V = 0.51 (Strong)
    Socio-economic level is associated with university-type preference.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether socio-economic level is associated with the preferred university type using chi-square. The
    result is very strong: chi-square(2) = 141.06, p < .001, Cramer V = 0.51. So the two variables are not
    independent -- socio-economic level markedly shapes university preference. This is an important dimension of
    opportunity inequality in education. For association between two categorical variables the chi-square test of
    independence is the standard method; Cramer V measures the strength of the association.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_learning_style.xlsx
  >> SCENARIO (narration):
    Suppose an educational psychologist wants to check whether learning styles
    are evenly distributed among students. We collected the dominant learning
    style of 200 students, classified as visual, auditory, kinesthetic, or
    enrolled. The question is whether the four learning-style categories occur
    in equal proportions, or whether some styles are clearly more common than
    others. Since we have a single categorical variable and want to compare
    observed category counts against an expected distribution, the Chi-Square
    Goodness-of-Fit test is appropriate.
  >> VARIABLE SELECTION:
    - Categorical variable: learning_style (visual/auditory/kinesthetic/enrolled)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 17.64   p < .001 ***   n = 200   k = 4 learning styles (expected: equal 1/k)
    H0 REJECTED

>> COMMENTARY (narration):
    We tested whether students' learning styles (visual, auditory, kinesthetic, reading-writing) are equally
    distributed with the chi-square goodness-of-fit test. The observed distribution deviates significantly from the
    equal expectation (chi-square(3) = 17.64, p < .001): some styles are more common than expected, others rare. So
    learning styles are not balanced in the student population -- suggesting instruction should be adapted to the
    dominant styles. The goodness-of-fit test is the right way to compare an observed category distribution against a
    theoretical expectation (equal or specific proportions).

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_method_achievement.xlsx
  >> SCENARIO (narration):
    Suppose we ran a small pilot study comparing two instruction methods and
    want to know whether passing a class depends on the method used. Only 28
    students took part, half taught with the traditional method and half with a
    mixed approach, and we recorded whether each one passed the class (yes or
    no). Our research question is whether method and class outcome are
    associated. Because this is a 2-by-2 table with a small sample where
    expected cell counts are low, Fisher's Exact Test is more reliable than the
    ordinary chi-square test.
  >> VARIABLE SELECTION:
    - Row variable: method (traditional/mixed)
    - Column variable: class_passed (no/yes)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher Exact p = 0.0377 *   Odds Ratio = 0.095   Phi = 0.43 (Strong)
    Instruction method is associated with class passing.   H0 REJECTED

>> COMMENTARY (narration):
    We tested the association between instruction method and class passing with Fisher's exact test -- because
    chi-square is unreliable for tables with few observations (small cells). The result is significant (p = 0.038),
    odds ratio 0.095: the odds of passing differ markedly between methods. Phi = 0.43 indicates a strong association.
    Fisher's exact test gives more accurate results than chi-square in small-sample pilot studies or with rare
    outcomes (small-cell 2x2 tables). It is often needed in educational pilot experiments.

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_student_achievement.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether a remediation program changed students'
    pass/fail status. We assessed the same 90 students twice: on an early test
    and a later test, recording for each whether they failed or succeeded. The
    research question is whether the proportion of students succeeding changed
    significantly from the first measurement to the second. Because we have
    paired, before-and-after binary outcomes on the same students, the McNemar
    Test is the correct choice for this matched dichotomous data.
  >> VARIABLE SELECTION:
    - First measurement (paired): on_test (failed/succeeded)
    - Second measurement (paired): last_test (failed/succeeded)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar chi-square = 0.89   p = 0.345   discordant: b = 17, c = 11   Cohen g = 0.21
    H0 NOT REJECTED (no significant change)

>> COMMENTARY (narration):
    We compared the same students' pre- and post-test pass status (pass/fail binary) with McNemar's test. This time
    the result is non-significant (chi-square = 0.89, p = 0.345): 17 students changed positively, 11 negatively --
    the balance is not tipped enough, so there is no significant shift in overall pass status. Non-significant results
    are informative too: this intervention did not change the binary pass status (it may have raised scores without
    pushing them over the threshold). McNemar looks only at the cells that changed (discordant cells); it is the
    right way to measure before/after categorical change.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_performance_teacher.xlsx
  >> SCENARIO (narration):
    Suppose two teachers independently rated the same students' performance and
    we want to know how well their judgments agree. Each of 70 students was
    graded by teacher A and teacher B using the same ordered categories (low,
    medium, good, high good). The question is not just how often they agree by
    chance, but how strong the genuine agreement is between the two raters.
    Cohen's Kappa is the right statistic because it measures inter-rater
    agreement for categorical ratings while correcting for agreement expected
    purely by chance.
  >> VARIABLE SELECTION:
    - Rater 1: teacher_a (low/medium/good/high good)
    - Rater 2: teacher_b (low/medium/good/high good)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's Kappa = 0.7144 (substantial/strong agreement)   p < .001 ***
    Two teachers' performance assessment

>> COMMENTARY (narration):
    We measured the agreement between two teachers independently assessing the same students with Cohen's Kappa.
    Kappa = 0.71 is in the "substantial/strong" range -- a chance-corrected measure of agreement, i.e. the real
    consistency after removing accidental agreement. In educational assessment this matters greatly: if teachers'
    grading is inconsistent, student scores vary by rater and unfairness results. Kappa = 0.71 shows the assessments
    are largely reliable and reproducible. Reporting inter-rater reliability is the quality-assurance step of
    rubric/performance assessment.

====================================================================================

#26  Cochran-Mantel-Haenszel Test
    file: 26_cmh_intervention_school.xlsx
  >> SCENARIO (narration):
    Suppose we evaluated an intervention across three different schools and want
    to know whether it improves student success while accounting for which
    school students attend. We have 360 students from school A, B, and C, each
    assigned to either the control or intervention group, with success recorded
    as yes or no. The research question is whether the intervention is
    associated with success after controlling for school. The Cochran-Mantel-
    Haenszel test is appropriate because it tests the group-outcome association
    across the strata defined by school.
  >> VARIABLE SELECTION:
    - Row variable: group (control/intervention)
    - Column variable: successful (no/yes)
    - Stratum variable: school (school-A/school-B/school-C)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi-square(1) = 13.89   p < .001 ***   Common OR (MH) = 2.36  (95% CI: 1.49 - 3.72)
    Breslow-Day chi-square(2) = 3.49, p = 0.17 (homogeneous OR)   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between intervention group and achievement while controlling for school (strata) with
    the CMH test. Even holding school constant the association is significant: common odds ratio 2.36 (p < .001) --
    stripped of the school confounder, the intervention group's odds of success are 2.4 times the control's. The
    Breslow-Day test is non-significant (p = 0.17), meaning the effect is consistent across all schools (homogeneous
    OR). CMH is the classic way -- very valuable in multi-centre education studies -- to control for a third variable
    (school, class) by stratification.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_sex_sound_university.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand how three categorical variables jointly relate
    to one another in a survey of 90 students. We recorded each student's sex
    (female/male), the noise level they study in (low/high), and whether they
    are an avid university reader (yes/no). Rather than testing a single
    pairwise association, we want to model the full pattern of associations and
    interactions among all three variables at once. Log-Linear Analysis is
    suited to this because it models the cell frequencies of the multi-way
    contingency table and reveals which main effects and interactions are needed
    to explain the data.
  >> VARIABLE SELECTION:
    - Factor 1: sex (female/male)
    - Factor 2: sound (low/high)
    - Factor 3: university_reader (no/yes)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Best model AIC = 46.72   Pearson chi-square = 0.0 (model fits the data perfectly)
    Factors: sex x socio-economic level x university preference

>> COMMENTARY (narration):
    We analysed the contingency structure formed jointly by three categorical variables -- sex, socio-economic level
    and university preference -- with a log-linear model. The selected model fits perfectly (Pearson chi-square = 0),
    with AIC = 46.72 as the most parsimonious model. Log-linear analysis lets us untangle interactions among more
    than two categorical variables (which are dependent, which interaction is significant) -- it is like ANOVA for
    categorical data. It is powerful for summarising multi-way demographic/preference relationships in education.

====================================================================================

#28  Crosstabulation
    file: 28_cross_age_political.xlsx
  >> SCENARIO (narration):
    Suppose we want to describe how political attitudes vary across age groups
    in our sample of 350 respondents. Each person is classified into an age
    group (18-30, 31-50, or 51+) and reports a political attitude (conservative,
    liberal, soloist, or uncertain). Here we are not testing a formal hypothesis
    so much as building a clear cross-tabulated picture of how attitudes are
    distributed within each age band, with counts and row percentages.
    Crosstabulation is exactly the tool for laying out the joint frequency
    distribution of these two categorical variables.
  >> VARIABLE SELECTION:
    - Row variable: age_group (18-30/31-50/51+)
    - Column variable: political_attitude (conservative/liberal/soloist/uncertain)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(6) = 15.81   p = 0.015 *   Cramer V = 0.15 (Weak-moderate)
    Age group is associated with attitude (weak).   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between age group and attitude/opinion with a crosstab and chi-square. The association
    is significant but weak (chi-square(6) = 15.81, p = 0.015, Cramer V = 0.15): age is related to attitude, but not
    strongly -- other factors (education, environment) are probably more determining. Reading significance together
    with effect size matters: in a large sample even a weak association can be significant. The crosstab is the most
    basic and readable way to see the co-occurrence of two categorical variables; Cramer V gives the true strength.

====================================================================================

#29  Multiple Response (MR) Frequency
    file: 29_mr_frequency_activities.xlsx
  >> SCENARIO (narration):
    Suppose we surveyed 250 students about the extracurricular activities they
    take part in, and each student could select several at once, things like
    sport, reading, computer, art, language course, theatre, and music. Because
    a single response field holds multiple comma-separated activities, we cannot
    simply count one value per student. Our goal is to find out how often each
    activity is chosen overall and what share of students engages in each one. A
    Multiple Response Frequency analysis is designed for exactly this kind of
    multi-select item, tallying each option across the whole sample.
  >> VARIABLE SELECTION:
    - Multiple-response set (comma-separated): activities
    - Response delimiter: comma

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)
    Extracurricular activities (sports, music, coding, volunteering...) selected as multi-choice

>> COMMENTARY (narration):
    Students were asked which extracurricular activities they take part in -- a "multiple response" question where
    several options can be ticked. All 250 students selected at least one activity, so the response rate is 100%.
    Multiple-response frequency analysis shows how many times and by what percentage of students each activity was
    chosen; the totals exceed 100% because everyone can tick several boxes. It is very useful in education for
    profiling students' interests and participation and for planning extracurricular activities.

====================================================================================

#30  Multiple Response Crosstab
    file: 30_mr_categorical_activity_sex.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether the kinds of extracurricular activities
    students choose differ between boys and girls. We surveyed 220 students who
    could each pick several activities from a multi-select list, and we also
    recorded each student's sex (male/female). The research question is how
    participation in each activity breaks down by sex. A Multiple Response
    Crosstab is the right approach because it cross-tabulates a multiple-
    response set against a single categorical grouping variable, showing counts
    and percentages of each activity within each sex.
  >> VARIABLE SELECTION:
    - Multiple-response set (comma-separated): activities
    - Crosstab (grouping) variable: sex (male/female)
    - Response delimiter: comma

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: sex
    Activities crossed by sex

>> COMMENTARY (narration):
    We cross-tabulated extracurricular activities (multiple-choice) by the student's sex. With data from 220 students
    we can see how each activity is distributed across sexes. The multiple-response crosstab answers "which group
    ticks which options more" -- e.g. girls lead in one activity, boys in another. This reveals the sex patterns in
    activity preference and guides inclusive activity planning.

====================================================================================

#31  MR x MR Crosstab
    file: 31_mr_mr_interest_career.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand how students' subject interests relate to the
    careers they aspire to. We surveyed 180 high-school students, and each one
    could tick more than one interest_area, choosing any combination of science,
    art, sport, and literature, while also selecting one or more aspirations in
    career_target among engineer, artist, doctor, and teacher. Because both
    variables are multiple-response sets rather than single picks, an ordinary
    crosstab would miscount students who chose several options. We use an MR by
    MR crosstab so MerQur can cross every interest category against every career
    category and report how often each pairing co-occurs, revealing, for
    example, whether science-oriented students disproportionately aim for
    engineering.
  >> VARIABLE SELECTION:
    - Multiple-response set 1 (rows): interest_area (science, art, sport, literature)
    - Multiple-response set 2 (columns): career_target (engineer, artist, doctor, teacher)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180
    Interest areas x career targets (teacher, doctor, engineer, artist...) cross-tabulated

>> COMMENTARY (narration):
    This is the most advanced multiple-response analysis: we cross two separate multi-choice questions against each
    other -- the student's interest areas and career targets. Among 180 students we can see which interests pair with
    which target careers: e.g. science-interested students target engineering/medicine, art-interested students
    target artist/teaching. This table reveals interest-career matches, providing powerful data for career guidance
    and counselling.

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q_course_sevme.xlsx
  >> SCENARIO (narration):
    Imagine we want to know whether students are equally likely to enjoy three
    different school subjects. We asked 150 students, for each of math, science,
    and language, simply whether they liked it, coded as yes or no. Since the
    same students answered all three subjects, these are three paired binary
    measurements on one group. We use Cochran's Q test because it is the right
    choice for comparing the proportion of yes responses across three or more
    related dichotomous conditions, telling us whether liking differs
    significantly among the subjects.
  >> VARIABLE SELECTION:
    - Related dichotomous variables (conditions): course = math / science / language, each scored liked (yes, no)
    - Subject identifier: student_id (auto-excluded)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000   (NO difference in liking proportion across courses)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the same students' "liking" (yes/no) proportions for different courses with Cochran's Q test -- the
    multi-occasion counterpart of repeated binary measures. The result is non-significant (Q = 0, p = 1.0): there is
    no difference in liking proportion across courses, students like the courses equally. A non-significant result is
    valuable too -- it says there is no distinction in course preference in this data. Cochran's Q is the right way to
    test change in repeated binary (yes/no) measurements on the same units; finding no expected difference here is a
    genuine result.

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5degisken.xlsx
  >> SCENARIO (narration):
    Suppose we want to map out how several academic and affective variables hang
    together for 180 students. We measured motivation, weekly study_hour, GPA,
    test anxiety, and overall school_satisfaction, and we would like to see
    which pairs move together and how strongly. Before building any predictive
    model, it is sensible to inspect all the bivariate relationships at once. We
    use a correlation analysis because it produces the full matrix of pairwise
    correlation coefficients with significance tests, so we can quickly spot,
    for instance, whether higher anxiety is associated with lower GPA.
  >> VARIABLE SELECTION:
    - Variables (correlation matrix): motivation
    - Variables (correlation matrix): study_hour
    - Variables (correlation matrix): GPA
    - Variables (correlation matrix): anxiety
    - Variables (correlation matrix): school_satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables (motivation, study hours, GPA, anxiety, school satisfaction)
    NO strong relationship (|r| >= 0.75) -- relationships are moderate/weak

>> COMMENTARY (narration):
    We computed all pairwise correlations among five key educational variables. No pair exceeds the |r| >= 0.75
    threshold -- so the relationships are moderate or weak. This is typical in education: motivation, study, anxiety
    are interrelated but none alone determines another; achievement is multi-factorial. The absence of strong
    relationships also means collinearity risk is low when building a regression -- it is safer to put the variables
    in the same model. The correlation matrix is the first step before modelling to see the structure among variables.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Analysis
    file: 34_bland_altman_IQ_test.xlsx
  >> SCENARIO (narration):
    Suppose a school psychology unit wants to know whether two intelligence
    tests can be used interchangeably. The same 60 students were assessed with
    both the WISC and the Stanford-Binet, giving us two IQ scores per child that
    should, in principle, measure the same construct. The question is not
    whether the scores correlate but whether they actually agree. We use a
    Bland-Altman analysis because it plots the difference between the two
    methods against their average and reports the mean bias and limits of
    agreement, which is exactly how method comparison should be evaluated.
  >> VARIABLE SELECTION:
    - Measurement 1 (M1): IQ_WISC
    - Measurement 2 (M2): IQ_Stanford_Binet

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    WISC mean = 100.18   |   Stanford-Binet mean = 97.98   |   Bias ~ +2.2 (methods agree largely)

>> COMMENTARY (narration):
    We applied two different IQ tests -- WISC and Stanford-Binet -- to the same students and compared them with
    Bland-Altman. This method does not look at correlation (high correlation does not guarantee agreement); it
    measures the difference between the two measurements. The means are very close (100.2 vs 98.0), the bias about 2
    points -- the two tests give almost the same result systematically and are practically interchangeable.
    Bland-Altman is the standard way to assess the agreement of two measurement tools (two tests, two raters,
    old/new method); it is ideal for showing test equivalence in educational measurement.

====================================================================================

#35  Effect Size
    file: 35_effect_size_material.xlsx
  >> SCENARIO (narration):
    Suppose we want to know not just whether instructional material matters, but
    how large its impact is. Sixty students learned with one of two materials, a
    printed book or an instructional video, and we recorded each student's
    test_score afterward. A simple significance test tells us whether the groups
    differ, but reviewers increasingly want a standardized measure of the
    practical magnitude. We use an effect size analysis because it quantifies
    the standardized difference between the book and video groups, for example
    Cohen's d, so we can communicate how meaningful the gap really is.
  >> VARIABLE SELECTION:
    - Grouping factor: material (book, video)
    - Outcome variable: test_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.60  (moderate effect)
    Test-score difference between two material groups

>> COMMENTARY (narration):
    We measured the test-score difference between two teaching materials not just as "is it significant" but "how
    large is it" -- as an effect size. Cohen's d = -0.60 is a moderate effect. P-values inflate with sample size, but
    effect size is sample-independent and gives the practical importance of the real difference. In education, "how
    much difference does this material/method make" is critical for resource-allocation decisions; a moderate effect
    of d = 0.60 says the intervention provides a notable but not miraculous benefit. Reporting effect size is standard
    in publications.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_bilissel_duyussal.xlsx
  >> SCENARIO (narration):
    Suppose we want to examine the relationship between students' cognitive
    abilities and their affective dispositions as two whole sets, rather than
    one variable at a time. For 150 students we have a cognitive battery of
    math, science, language, reading_comprehension, and logic, alongside an
    affective battery of motivation, attitude, satisfaction, and anxiety. The
    research question is how these two domains co-vary jointly. We use canonical
    correlation analysis because it finds the linear combinations of each set
    that are maximally correlated, summarizing the cognitive-affective link in a
    few interpretable canonical dimensions.
  >> VARIABLE SELECTION:
    - Set X (cognitive): math
    - Set X (cognitive): science
    - Set X (cognitive): language
    - Set X (cognitive): reading_comprehension
    - Set X (cognitive): logic
    - Set Y (affective): motivation
    - Set Y (affective): attitude
    - Set Y (affective): satisfaction
    - Set Y (affective): anxiety

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.977, chi-square(20) = 1380.35, p < .001   |   CC2: r = 0.959, p < .001
    X set: math/science/language/reading/logic (cognitive)  |  Y set: motivation/attitude/satisfaction/anxiety (affective)

>> COMMENTARY (narration):
    We related two multivariate sets -- cognitive skills (math, science, language, reading, logic) and affective
    variables (motivation, attitude, satisfaction, anxiety) -- with canonical correlation. The first canonical axis
    is extremely strong (r = 0.98, p < .001), the second also significant (r = 0.96): the cognitive and affective
    domains are very tightly linked. So academic skills and motivation/attitude move together -- high cognitive
    ability coincides with high motivation and low anxiety. CCA is the most direct answer to "how do two blocks of
    variables co-vary"; it is a powerful method for studying cognitive-affective interplay in educational psychology.

====================================================================================

#37  Correspondence Analysis (CA)
    file: 37_ca_sound_occupation.xlsx
  >> SCENARIO (narration):
    Suppose we want to explore how the noise level students study in relates to
    the occupations they prefer. Among 470 students we recorded their typical
    study sound environment as low, medium, or high, and their
    occupation_preference among engineer, teacher, civil servant, artist,
    doctor, and academic. With two categorical variables and many cells, raw
    counts are hard to interpret. We use correspondence analysis because it
    turns the contingency table into a two-dimensional map, visually showing
    which sound levels sit close to which career preferences.
  >> VARIABLE SELECTION:
    - Row variable: sound (low, medium, high)
    - Column variable: occupation_preference (engineer, teacher, civil_servant, artist, doctor, academic)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.608   Number of dimensions = 2
    Socio-economic level x occupation preference contingency table

>> COMMENTARY (narration):
    We projected the contingency table of socio-economic level by occupation preference into a two-dimensional map
    with correspondence analysis. Total inertia 0.608 indicates a strong association; the two dimensions visualise
    most of that structure. Level-occupation pairs that fall close on the map indicate that socio-economic group
    leans toward that occupation. This lets us read at a glance how socio-economic background shapes occupational
    expectations in education. It is the most powerful way to visualise categorical relationships.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_18_item.xlsx
  >> SCENARIO (narration):
    Suppose we administered an 18-item psychological scale to 180 students and
    we want to confirm whether the items naturally group into the three intended
    constructs. The items are six self-efficacy items, six motivation items, and
    six anxiety items, all on numeric scales. Rather than assuming the
    structure, we want the data to reveal which items cluster together. We use
    variable clustering because it groups the 18 items into clusters of mutually
    correlated variables, letting us verify that the self-efficacy, motivation,
    and anxiety items really form distinct families.
  >> VARIABLE SELECTION:
    - Variables to cluster (items): self_efficacy_m1, self_efficacy_m2, self_efficacy_m3, self_efficacy_m4, self_efficacy_m5, self_efficacy_m6
    - Variables to cluster (items): motivation_m1, motivation_m2, motivation_m3, motivation_m4, motivation_m5, motivation_m6
    - Variables to cluster (items): anxiety_m1, anxiety_m2, anxiety_m3, anxiety_m4, anxiety_m5, anxiety_m6

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: self-efficacy / motivation / anxiety)
    A representative item was selected from each cluster

>> COMMENTARY (narration):
    We clustered 18 scale items by similarity into three groups. VarClus clusters items, not observations: it shows
    which items measure the same "construct". The three clusters formed exactly as expected: self-efficacy items,
    motivation items and anxiety items clustered separately. This confirms the scale really measures three
    sub-dimensions and lets us pick one representative per group to reduce the number of items. In scale development
    and item reduction it is extremely practical for weeding out redundancy.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_GPA.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict students' grade point average from a set of
    background and behavioral factors. For 200 students we have weekly
    study_hour, motivation, parent_education_year, and absenteeism_day, and we
    want to know which of these contribute to GPA and by how much. The outcome
    is continuous and we have several numeric predictors to weigh
    simultaneously. We use multiple linear regression because it estimates the
    unique effect of each predictor on GPA while holding the others constant,
    and tells us how much variance the model explains overall.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Predictors: study_hour
    - Predictors: motivation
    - Predictors: parent_education_year
    - Predictors: absenteeism_day

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.580   Adjusted R2 = 0.572
    GPA ~ study hours + motivation + parent education + absenteeism

>> COMMENTARY (narration):
    We built a multiple linear regression predicting GPA from four variables -- study hours, motivation, parent
    education, absenteeism. The model is moderately strong: together they explain 58% of GPA variance (R2 = 0.58).
    This is a good explanatory rate in education -- achievement is multi-factorial and 58% captures a meaningful
    portion; the rest is unmeasured factors (ability, teacher, home environment) and individual differences. Such a
    model shows which variable contributes most to achievement and helps identify at-risk students early. Multiple
    regression is the basic tool for explaining a continuous outcome with several drivers.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_university_achievement.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict whether a student will succeed at the university
    entrance level from their earlier profile. For 300 students we recorded GPA,
    study_hour, motivation, and study sound_level, and the outcome
    university_achievement is binary, success or not. Because the response is
    dichotomous rather than continuous, linear regression is inappropriate. We
    use logistic regression because it models the probability of university
    success as a function of these predictors and reports odds ratios, showing
    how each factor shifts a student's likelihood of passing.
  >> VARIABLE SELECTION:
    - Target (binary outcome): university_achievement (0/1)
    - Predictors: GPA
    - Predictors: study_hour
    - Predictors: motivation
    - Predictors: sound_level

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: Logistic Regression   Pseudo R2 = 0.10
    University achievement (1/0) ~ GPA + study hours + motivation + socio-economic level

>> COMMENTARY (narration):
    We built a logistic regression predicting a student's probability of getting into university from GPA, study,
    motivation and socio-economic level. Because the outcome is binary (admitted/not), logistic regression is the
    right choice. Pseudo R2 = 0.10 shows the model partly explains the outcome -- the coefficients generally say the
    odds of admission rise as GPA and study increase. In education, logistic regression is the most common way to
    predict an entry/outcome (university, graduation, passing); it gives each factor's effect on the odds in an
    interpretable way.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Poisson / Negative Binomial Regression
    file: 41_poisson_discipline_count.xlsx
  >> SCENARIO (narration):
    Suppose we are studying disciplinary incidents in secondary schools. For 180
    students we recorded their age, their academic motivation, and the count of
    disciplinary actions they received over the year (annual_discipline_count).
    Because the outcome is a count of events rather than a continuous measure,
    we want to model how age and motivation relate to that count. We use Poisson
    or Negative Binomial regression here since the dependent variable is a non-
    negative integer count, and Negative Binomial is the right choice when the
    counts show overdispersion.
  >> VARIABLE SELECTION:
    - Target (count DV): annual_discipline_count
    - Predictor: age
    - Predictor: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 187.43   Deviance = 117.59
    Annual discipline count ~ age + motivation   (Poisson)

>> COMMENTARY (narration):
    We built a Poisson regression predicting students' annual discipline-incident count -- a count variable -- from
    age and motivation. Count data (0,1,2,...) are not normally distributed, so we use Poisson rather than linear
    regression. The model typically shows discipline incidents falling as motivation rises. Deviance and AIC assess
    model fit. Count regression is the correct way to relate count outcomes in education (discipline incidents,
    absences, applications) to predictors; it is far more appropriate than a linear model.

====================================================================================

#42  Multinomial Logistic Regression
    file: 42_multinomial_occupation_preference.xlsx
  >> SCENARIO (narration):
    Imagine we want to understand what drives students' career aspirations. For
    250 students we have their math_score and their art_score, and each student
    named an occupation_preference that is one of doctor, artist, or engineer.
    The research question is whether academic strengths in math versus art
    predict which career a student leans toward. We use Multinomial Logistic
    Regression because the outcome is an unordered categorical variable with
    more than two levels.
  >> VARIABLE SELECTION:
    - Target (nominal DV): occupation_preference (doctor, artist, engineer)
    - Predictor: math_score
    - Predictor: art_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 298.45   Reference class: 'artist'
    Occupation preference ~ math score + art score

>> COMMENTARY (narration):
    We modelled students' occupation preference (several categories) from math and art scores with multinomial
    logistic regression. Because there are more than two unordered categories, this method fits: 'artist' is the
    reference, and a separate equation is built for each other occupation. Coefficients read as "how many times the
    odds of becoming an engineer rather than an artist change as the math score rises". This shows how a student's
    ability profile shapes occupational orientation -- a valuable tool in career guidance and counselling.

====================================================================================

#43  Ordinal Logistic Regression
    file: 43_ordinal_motivation_level.xlsx
  >> SCENARIO (narration):
    Here we examine what shapes student motivation. For 200 students we measured
    perceived teacher_support and school_opportunity, and we classified each
    student's motivation_level into ordered categories from low to high. We want
    to know whether more support and more opportunity move students up the
    motivation ladder. We use Ordinal Logistic Regression because the outcome is
    an ordered categorical variable, so the natural ranking among the motivation
    levels is respected.
  >> VARIABLE SELECTION:
    - Target (ordinal DV): motivation_level (low, medium, high, high high)
    - Predictor: teacher_support
    - Predictor: school_opportunity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 441.39
    Motivation level (low<medium<high) ~ teacher support + school opportunity

>> COMMENTARY (narration):
    We built an ordinal logistic regression predicting motivation level (low, medium, high -- ordered) from teacher
    support and school opportunity. These classes are ordered -- there is a natural ranking -- and the ordinal model
    uses that order, making it more powerful and interpretable than multinomial. The coefficients give a student's
    tendency to move to a higher motivation level as teacher support and school opportunity increase. For ordered
    categorical outcomes (motivation level, Likert, achievement tier), ordinal logistic is the right method.

====================================================================================

#44  PLS Regression
    file: 44_pls_behavior_achievement.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict academic achievement from a battery of twelve
    correlated classroom behavior indicators (behavior_01 through behavior_12)
    measured on 200 students. Because these behavior items are highly
    intercorrelated and there are many of them relative to a single outcome,
    ordinary regression would be unstable. We use PLS Regression because it
    extracts latent components from the predictor block that best explain
    academic_achievement, handling the multicollinearity among the behavior
    items.
  >> VARIABLE SELECTION:
    - Predictors (X block): behavior_01, behavior_02, behavior_03, behavior_04, behavior_05, behavior_06, behavior_07, behavior_08, behavior_09, behavior_10, behavior_11, behavior_12
    - Target (Y): academic_achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2_train = 0.908   R2_CV (5-fold) = 0.903
    Academic achievement ~ 12 behavior items

>> COMMENTARY (narration):
    We predicted academic achievement from 12 inter-correlated behavior items with PLS (Partial Least Squares)
    regression. When the behavior items are correlated, ordinary regression becomes unstable; PLS solves this by
    reducing the items to a few orthogonal components. The model is very strong both in training (R2 = 0.91) and
    cross-validation (R2_CV = 0.90) -- no overfitting, excellent generalisation. When there are many correlated
    items/behaviors (observation scales, survey batteries), PLS is the preferred method that keeps both predictive
    power and interpretability.

====================================================================================

#45  Probit Regression
    file: 45_probit_study_achievement.xlsx
  >> SCENARIO (narration):
    Imagine we are looking at exam outcomes for 180 students. We recorded each
    student's weekly_study hours and whether they passed the exam, coded as
    passed (0 = fail, 1 = pass). We want to estimate how the probability of
    passing changes with study time. We use Probit Regression because the
    outcome is binary and we model the pass probability through the cumulative
    normal distribution, an alternative to logistic regression for a dichotomous
    response.
  >> VARIABLE SELECTION:
    - Target (binary DV): passed
    - Predictor: weekly_study

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 127.62   Pseudo R2 (McFadden) = 0.494
    Passing (1/0) ~ weekly study

>> COMMENTARY (narration):
    We predicted whether a student passes from weekly study hours with probit regression. Probit models a binary
    outcome like logistic but assumes a normal-distribution curve (cumulative normal). Pseudo R2 = 0.49 is quite
    strong: study hours largely explain the probability of passing -- a student who studies more has a markedly
    higher chance of passing. Probit and logistic usually give similar results; the choice is often tradition and
    interpretive preference. It suits predicting achievement/passing probability from a continuous predictor in
    education.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_education_expenditure.xlsx
  >> SCENARIO (narration):
    Suppose we study household spending on education. For 200 families we have
    family_income_TL, child_count, and education_expenditure_TL. Many families
    report zero education spending, so the outcome is censored at zero and a
    standard linear model would be biased. We use Tobit Regression because the
    dependent variable is left-censored at zero, letting us correctly model how
    income and number of children relate to education expenditure.
  >> VARIABLE SELECTION:
    - Target (censored DV): education_expenditure_TL
    - Predictor: family_income_TL
    - Predictor: child_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 414.25   Tobit (left-censored: expenditure floored)
    Education expenditure ~ family income + child count

>> COMMENTARY (narration):
    We modelled families' education expenditure from income and number of children. The problem: some families spend
    nothing on education (zero) or expenditure piles up below a floor value -- this is censored data, and ordinary
    regression gives biased results. Tobit regression correctly models this left-censored structure (AIC = 414.25).
    Typically expenditure rises with income and falls with more children. The statistically correct solution for
    working with censored/piled-up data (expenditure, duration, dosage) is the Tobit model.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_study_achievement.xlsx
  >> SCENARIO (narration):
    Here we model test performance for 100 students using their weekly
    study_hour and their age. Rather than relying only on point estimates, we
    want full probability distributions for the regression coefficients and to
    incorporate uncertainty directly. We use Bayesian Linear Regression because
    it yields posterior distributions and credible intervals for how study hours
    and age predict the test_score, which is a continuous outcome.
  >> VARIABLE SELECTION:
    - Target (DV): test_score
    - Predictor: study_hour
    - Predictor: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Conjugate Normal-Inverse-Gamma prior   posterior sigma2 = 23.72
    Test score ~ study hours + age

>> COMMENTARY (narration):
    We used Bayesian linear regression to predict test score from study hours and age. The Bayesian advantage: the
    result is not a single p-value but a probability distribution (posterior) for the parameters -- we get a credible
    interval for each coefficient and a "probability the effect is positive". The posterior error variance is sigma2
    = 23.72. The Bayesian method honestly expresses uncertainty and can incorporate prior knowledge. It is valuable
    in educational research with small samples or where information must be carried over from previous studies.

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear_language_age.xlsx
  >> SCENARIO (narration):
    Imagine we are tracking early language development. For 80 children we
    recorded age_year and their language_word_count. Vocabulary growth is not a
    straight line; it rises steeply and then levels off, so a linear fit is
    inappropriate. We use Nonlinear Regression because the relationship between
    age and word count follows a curved, saturating growth pattern that we want
    to capture with an appropriate nonlinear function.
  >> VARIABLE SELECTION:
    - Target (DV): language_word_count
    - Predictor: age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: y = K / (1 + exp(-r·(x-x0)))  (logistic growth)   R2 = 0.943
    Language (word count) ~ age

>> COMMENTARY (narration):
    We modelled children's vocabulary growth by age. Language acquisition is not linear: it starts slow, rises fast,
    then saturates toward a ceiling (K) -- a classic S-curve (logistic growth). We modelled this with a logistic
    function and the fit is excellent (R2 = 0.94). The curve shows at which age development accelerates and when it
    approaches the ceiling. Nonlinear regression is the way to correctly model growth/learning curves -- language
    acquisition, skill gain, forgetting curves; forcing a straight line would misrepresent the nature of development.

====================================================================================

#49  Ridge Regression
    file: 49_ridge_item_motivation.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict student motivation from a 15-item questionnaire
    (item_p01 through item_p15) administered to 200 students. The items are
    strongly correlated with one another, which makes ordinary least squares
    coefficients unstable. We use Ridge Regression because its L2 penalty
    shrinks the coefficients and stabilizes the model under multicollinearity,
    giving a more reliable prediction of motivation from the item set.
  >> VARIABLE SELECTION:
    - Predictors: item_p01, item_p02, item_p03, item_p04, item_p05, item_p06, item_p07, item_p08, item_p09, item_p10, item_p11, item_p12, item_p13, item_p14, item_p15
    - Target (DV): motivation

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Ridge (alpha = 1.0): R2 = 0.554   Adj. R2 = 0.518   n = 200
    Motivation ~ 15 scale items (mutually correlated)

>> COMMENTARY (narration):
    Predicting motivation from 15 highly correlated scale items, ordinary regression becomes unstable (collinearity
    inflates coefficients). Ridge regression solves this by shrinking all coefficients -- it does not zero them but
    limits their magnitude, keeping prediction stable (R2 = 0.55). Ridge is ideal when "I want to keep all items but
    there is collinearity". For naturally correlated variable sets like scale items, it prevents overfitting and
    provides stable estimates.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso_30_pred_5_active.xlsx
  >> SCENARIO (narration):
    Imagine we have 30 candidate predictors (x01 through x30) measured on 250
    students, and we want to model academic_achievement. We suspect that only a
    handful of these variables actually matter and we want the model to perform
    automatic variable selection. We use Lasso Regression because its L1 penalty
    drives the coefficients of irrelevant predictors exactly to zero, leaving a
    sparse, interpretable subset that predicts achievement.
  >> VARIABLE SELECTION:
    - Predictors: x01, x02, x03, x04, x05, x06, x07, x08, x09, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24, x25, x26, x27, x28, x29, x30
    - Target (DV): academic_achievement

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Lasso (alpha = 0.1): R2 = 0.722   Adj. R2 = 0.684   n = 250
    16 of 30 predictor coefficients shrunk to zero (automatic feature selection)

>> COMMENTARY (narration):
    In this data a few real predictors were mixed with many irrelevant variables. The strength of Lasso shows here:
    the penalty term drives irrelevant variables' coefficients exactly to zero -- automatic variable selection. 16 of
    30 predictors were eliminated, and the model stayed strong at R2 = 0.72. Lasso is extremely useful when you want
    to pick the true drivers from many candidates -- in high-dimensional educational/psychometric data. While Ridge
    shrinks all variables, Lasso eliminates the irrelevant ones entirely.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_parent_motivation.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand how parental support shapes students' academic
    achievement. We have data from 200 students, each with a score for perceived
    parent_support, their academic motivation, and their academic_achievement.
    Our hypothesis is that parental support doesn't act directly on achievement
    so much as it works through motivation: supportive parents raise students'
    motivation, and that motivation in turn drives achievement. To test this
    indirect pathway we use mediation analysis, treating parent_support as the
    predictor, motivation as the mediator, and academic_achievement as the
    outcome, which lets us decompose the total effect into direct and indirect
    components.
  >> VARIABLE SELECTION:
    - X (predictor): parent_support
    - M (mediator): motivation
    - Y (outcome): academic_achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 5.55   p < .001   95% CI (3.83 , 7.30)   Significant
    Path: parent support -> motivation -> academic achievement

>> COMMENTARY (narration):
    We tested whether parent support's effect on achievement passes through motivation with mediation analysis. The
    indirect effect is significant (5.55, p < .001, CI excludes zero): parent support raises motivation, and higher
    motivation raises achievement. So parent support affects achievement not directly but "through" motivation. This
    opens up one of the core questions of educational psychology -- the mechanism by which family influence works.
    Mediation analysis reveals the intermediate mechanism in "why X affects Y" and points to intervention targets.

====================================================================================

#52  Path Analysis
    file: 52_path_oz_yeterlik_achievement.xlsx
  >> SCENARIO (narration):
    Here we are interested in the psychological mechanism behind academic
    success in a sample of 220 students. We measured each student's
    self_efficacy, their motivation, the effort they put into their studies, and
    their final achievement. Theory suggests these variables form a chain: self-
    efficacy fuels motivation, motivation drives effort, and effort produces
    achievement, with some direct links as well. Path analysis is the right tool
    because it lets us specify and test this entire system of directed
    relationships simultaneously rather than running separate regressions.
  >> VARIABLE SELECTION:
    - Exogenous variable: self_efficacy
    - Mediators: motivation
    - Mediators: effort
    - Endogenous (outcome): achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.930   RMSEA = 0.193   (poor fit -- model should be revised)
    Model: self-efficacy/motivation/effort -> achievement; self-efficacy -> motivation

>> COMMENTARY (narration):
    Path analysis is extended regression that models several cause-effect relationships at once. Here we tested that
    self-efficacy, motivation and effort affect achievement, and that self-efficacy feeds motivation. CFI = 0.93 is
    acceptable but RMSEA = 0.19 is high -- the hypothesised model does not fully match the data; some paths may be
    missing or extra. This is informative too: it says the theoretical model should be revised. Path analysis is a
    powerful way to separate direct and indirect effects and test theoretical models; both good and poor fit guide us
    in improving the model.

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_student_term_GPA.xlsx
  >> SCENARIO (narration):
    Imagine we tracked 200 students' GPA across four academic terms and want to
    know how grades evolve over time. Because each student contributes several
    repeated GPA measurements, the observations are nested within students and
    are not independent. A linear mixed model is appropriate here: we model GPA
    as the outcome with term as a fixed effect to capture the average
    trajectory, while a random intercept for student_id absorbs the stable
    between-student differences in overall performance.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Fixed effect (factor): term (term 1/term 2/term 3/term 4)
    - Grouping / random effect (subject): student_id

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group variance (random intercept) = 0.248   ICC = 0.942
    GPA ~ term  +  (1 | student)

>> COMMENTARY (narration):
    Modelling the same students' GPA across terms, the within-student repeated measures are not independent. The
    linear mixed model solves this by adding student as a random effect. ICC = 0.94 is very high: 94% of GPA variance
    comes from between-student differences; within-term change is small. So students are consistent within
    themselves, and the main difference is between individuals. LMM is the correct, indispensable method for
    nested/repeated-measures data (student>term, school>student) -- it models each student's own trajectory and
    resolves the independence violation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation_missing_student.xlsx
  >> SCENARIO (narration):
    Suppose we collected GPA, study_hour, motivation, anxiety, and
    parent_education from 150 students, but as often happens in survey research
    some values are missing across these variables. Simply deleting incomplete
    cases would shrink our sample and could bias the results if the data are not
    missing completely at random. Multiple imputation is the principled
    solution: it generates several plausible complete datasets by drawing
    missing values from their predictive distribution, so we can analyze the
    full sample while properly reflecting the uncertainty introduced by the
    missing data.
  >> VARIABLE SELECTION:
    - Variables to impute: GPA
    - Variables to impute: study_hour
    - Variables to impute: motivation
    - Variables to impute: anxiety
    - Variables to impute: parent_education

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (number of imputations) = 5
    Missing GPA values estimated from study/motivation/anxiety/parent education

>> COMMENTARY (narration):
    Missing observations are inevitable in educational data (absent students, blank survey items); deleting them both
    loses data and biases results. Multiple imputation estimates the missing GPA values from other variables,
    creating 5 separate completed datasets, runs the analysis in each and combines the results -- thereby also
    accounting for estimation uncertainty. Under the MAR (missing at random) assumption this is the modern
    missing-data standard. In longitudinal educational studies it is the soundest way to cope with lost student data.

====================================================================================

#55  Generalized Estimating Equations (GEE)
    file: 55_gee_motivation_group.xlsx
  >> SCENARIO (narration):
    Here we ran a longitudinal intervention study where 300 student records
    track motivation across repeated visits, with each student assigned to
    either a control or an intervention condition. Because the same students are
    measured at several visits, their motivation scores are correlated over
    time. Generalized estimating equations let us model the population-averaged
    effect of group and visit on motivation while accounting for this within-
    subject correlation through a working correlation structure, giving us
    robust estimates of how the intervention shifts motivation on average.
  >> VARIABLE SELECTION:
    - Dependent variable: motivation
    - Subject (cluster id): student_id
    - Within-subject (time): visit
    - Predictor (group): group (control/intervention)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coefficient = 0.020   p = 0.24 (non-significant)   QIC = 305.24
    Motivation ~ visit  (repeated measures, clustered within student)

>> COMMENTARY (narration):
    We took repeated (multi-visit) motivation measurements from the same students -- one student's measurements are
    dependent. GEE estimates the relationship at the "population average" level in such clustered data, accounting
    for the correlation structure. The visit effect is non-significant (p = 0.24): motivation did not change
    systematically over time. A non-significant result is information too -- the intervention/time may not have
    changed motivation. GEE is an alternative to LMM: rather than modelling random effects, it corrects the
    correlation among observations and gives the marginal (average) effect.

====================================================================================

#56  Generalized Linear Mixed Model (GLMM)
    file: 56_glmm_discipline_count.xlsx
  >> SCENARIO (narration):
    Suppose we are studying disciplinary incidents in 160 student-month records,
    counting the discipline_count for each student over several months, and
    noting whether a behavioral intervention_present was active. The outcome is
    a count that varies repeatedly within each student, so it is both non-normal
    and clustered. A generalized linear mixed model fits naturally: we model
    discipline_count with a Poisson family, include month and
    intervention_present as fixed effects, and add a random intercept for
    student_id to capture each student's baseline tendency toward discipline
    incidents.
  >> VARIABLE SELECTION:
    - Dependent variable (count): discipline_count
    - Fixed effect: intervention_present
    - Fixed effect: month
    - Grouping / random effect (subject): student_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    month: coef = 0.157, p = 0.006 *   |   intervention_present: coef = -0.510, p < .001 ***
    Poisson GLMM: discipline count ~ month + intervention  +  (1 | student)

>> COMMENTARY (narration):
    We modelled discipline counts (Poisson) with both month/intervention fixed effects and a student random effect --
    a GLMM: mixed model + count distribution. The intervention effect is significant and negative (coef -0.51,
    p < .001): the intervention markedly reduces discipline incidents. The month effect is also significant (an
    upward tendency). GLMM combines the strengths of LMM (which assumes normality) and GLM (which ignores clustering)
    for "repeated/nested count or binary data". It is the correct framework for repeated count data per student --
    monitoring a behavioral intervention.

====================================================================================

#57  Regularized Regression (Elastic Net)
    file: 57_elasticnet_achievement_index.xlsx
  >> SCENARIO (narration):
    Imagine we want to predict a composite achievement_index for 220 students
    from 40 candidate predictors (x01 through x40). With this many features,
    ordinary regression risks overfitting and the predictors are likely
    correlated with one another. Regularized regression using Elastic Net is
    ideal here because it blends ridge and lasso penalties: it shrinks
    coefficients to control overfitting, handles correlated predictors
    gracefully by grouping them, and can drive irrelevant coefficients to zero
    for a more interpretable, generalizable model.
  >> VARIABLE SELECTION:
    - Target: achievement_index
    - Predictors: x01–x40 (all 40 numeric predictors)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.087   R2 (train) = 0.501
    Achievement index ~ 40 predictors

>> COMMENTARY (narration):
    Elastic Net is a hybrid of Ridge and Lasso: it both shrinks coefficients (like Ridge) and zeros out the
    irrelevant ones (like Lasso). Here we predicted an achievement index from 40 predictors; Elastic Net keeps
    correlated variable groups together while eliminating the irrelevant ones. The model gives R2 = 0.50. When there
    are many collinear predictors -- item batteries, behavior scales, demographics -- Elastic Net is a balanced
    choice offering both selection and stability. It softens Lasso's "pick one at random from a correlated group"
    problem.

====================================================================================

#58  Robust Regression
    file: 58_robust_study_achievement.xlsx
  >> SCENARIO (narration):
    Suppose we are modeling how weekly study_hour predicts a student's
    test_score in a sample of 100 students, but a few students show unusual
    patterns, such as very high scores with little study or the reverse. These
    outliers can heavily distort an ordinary least squares line. Robust
    regression is the appropriate choice because it down-weights the influence
    of extreme observations, giving us a regression of test_score on study_hour
    that reflects the bulk of the students rather than being dragged around by a
    handful of atypical cases.
  >> VARIABLE SELECTION:
    - Dependent variable: test_score
    - Predictor: study_hour

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust intercept = 49.06   vs   OLS intercept = 42.20   (outliers distorted OLS)
    Test score ~ study hours

>> COMMENTARY (narration):
    This data had a few outliers -- e.g. students who studied a lot but scored low, or studied little but scored
    high. Ordinary regression (OLS) takes outliers fully into account and is pulled toward them: the intercept is
    42.20 in OLS but 49.06 in the robust model; the large gap means outliers distorted OLS. Robust regression
    (Huber-T) gives less weight to outlying observations to preserve the true trend. In educational research, where
    outlier student data are inevitable, robust regression gives a more reliable slope estimate than OLS.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_GPA.xlsx
  >> SCENARIO (narration):
    Here we want to know how study_hour and motivation relate to GPA across 200
    students, but not just at the average. We suspect these predictors matter
    differently for struggling students versus high achievers. Quantile
    regression is the right approach because instead of modeling only the
    conditional mean, it estimates the effect of study_hour and motivation at
    different quantiles of the GPA distribution, revealing whether, say,
    studying more helps low-GPA students more than it helps those already near
    the top.
  >> VARIABLE SELECTION:
    - Dependent variable: GPA
    - Predictors: study_hour
    - Predictors: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    q = 0.10 slope = 0.28   |   q = 0.50 (median)   |   q = 0.90  (varying effect)
    GPA ~ study hours + motivation

>> COMMENTARY (narration):
    Ordinary regression models only the mean; but study's effect on GPA may differ for low- and high-achieving
    students. Quantile regression models the lower (q=0.10, low GPA), middle (q=0.50) and upper (q=0.90, high GPA)
    quantiles separately. A changing slope across quantiles answers "does study help every student equally" -- for
    instance study may help low achievers more. In education, when the different parts of the distribution (weakest/
    strongest students) matter rather than the average, quantile regression reveals what mean-based models miss.

====================================================================================

#60  ROC Curve Analysis
    file: 60_roc_composite_university_achievement.xlsx
  >> SCENARIO (narration):
    Suppose we built a composite_achievement_score for 250 students and want to
    know how well it predicts whether a student is later university_successful,
    a binary outcome coded 0/1. We need to evaluate the score's discriminative
    ability across all possible cut-off thresholds rather than picking one
    arbitrarily. ROC curve analysis is exactly suited to this: it plots
    sensitivity against one-minus-specificity over every threshold and
    summarizes overall accuracy with the area under the curve, while also
    helping us identify an optimal cut-off for classifying students.
  >> VARIABLE SELECTION:
    - Classifier / test score: composite_achievement_score
    - Binary outcome (event): university_successful

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.933 (excellent discriminating power)   Youden optimum threshold: 0.54 -> Sensitivity 0.90
    n = 250 (positive 143, negative 107)   university achievement classifier

>> COMMENTARY (narration):
    We measured how well a composite achievement score predicts university success with the ROC curve. AUC = 0.93 is
    "excellent": the score largely separates admitted from non-admitted students correctly. The Youden index gives
    the best cut-off (threshold), optimising sensitivity and specificity; here sensitivity is 0.90 at a 0.54
    threshold. ROC/AUC is the standard for reporting a diagnostic/classification model's discriminating power; in
    education it is ideal for evaluating selection exams, risk-screening tools and prediction models.

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss_class_pass.xlsx
  >> SCENARIO (narration):
    Imagine we have built a model to predict whether students will pass their
    class, and we want to evaluate how good that prediction really is. For 200
    students we recorded a continuous prediction_score from the model along with
    the actual class_pass outcome, where 1 means the student passed and 0 means
    they failed. The True Skill Statistic is ideal here because it summarizes
    both sensitivity and specificity into a single threshold-independent measure
    of predictive skill, telling us how much better our model performs than
    random guessing. It is especially useful in education when pass and fail
    rates are unbalanced.
  >> VARIABLE SELECTION:
    - Predicted score: prediction_score
    - Observed outcome (binary): class_pass (1 = pass, 0 = fail)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.481 (good)   Sensitivity 0.85   Specificity 0.63   n = 200
    Class-passing prediction

>> COMMENTARY (narration):
    TSS (True Skill Statistic) measures how much better than chance a binary classification model is: sensitivity +
    specificity - 1. The model predicting class passing has TSS = 0.48 -- "good": sensitivity 0.85 (high at catching
    passers), specificity 0.63 (moderate at separating those who fail). TSS's advantage is insensitivity to
    prevalence (how common passing is), so it is reliable for imbalanced classes. It summarises the true skill of
    achievement/risk prediction models in education, getting past the bias that accuracy hides.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity_achievement_3sinif.xlsx
  >> SCENARIO (narration):
    Suppose an automated grading system assigns each student a predicted
    achievement label, and we want to check how accurate it is against the true
    labels. For 200 assessment units we have the actual_label, the model's
    prediction_label, and a continuous prediction_score behind each decision.
    Confusion Matrix Metrics let us compute accuracy, precision, recall, and F1
    from the cross-tabulation of predicted versus actual classes. This is the
    natural choice when we have a categorical classifier and want a full
    breakdown of where it gets things right and where it confuses one class for
    another.
  >> VARIABLE SELECTION:
    - Actual class: actual_label
    - Predicted class: prediction_label
    - Predicted probability/score: prediction_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.93   F1 = 0.918
    Achievement-class prediction -- actual vs predicted label

>> COMMENTARY (narration):
    We broke down a classifier's achievement-class predictions with a confusion matrix. With accuracy 93% and F1 =
    0.92 the model is very successful -- both the overall correct rate is high and the precision/recall balance (F1)
    is good. The confusion matrix reveals what the "accuracy" number hides -- where each class is misclassified,
    which classes are confused. The F1 score is more informative than accuracy especially with imbalanced classes.
    It is the indispensable tool for evaluating the real performance of automated grading/classification models in
    education.

====================================================================================

#63  Random Forest
    file: 63_rf_occupation_preference.xlsx
  >> SCENARIO (narration):
    Let's say we want to predict which occupation students will prefer based on
    their academic profile. For 350 students we have scores in math, science,
    language, and reading_comprehension, along with their overall GPA, and the
    target occupation_preference takes one of four values: teacher, engineer,
    doctor, or artist. Random Forest is a great fit because it handles several
    continuous predictors at once, captures nonlinear relationships and
    interactions automatically, and gives us a ranking of which subjects matter
    most for predicting career preference, all while resisting overfitting.
  >> VARIABLE SELECTION:
    - Target (class): occupation_preference (teacher / engineer / doctor / artist)
    - Predictors: math
    - Predictors: science
    - Predictors: language
    - Predictors: reading_comprehension
    - Predictors: GPA

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.814   (occupation preference classification)
    Predicted from 5 academic variables

>> COMMENTARY (narration):
    We predicted students' occupation preference from five academic variables (math, science, language, reading, GPA)
    with Random Forest -- the vote of hundreds of decision trees. Accuracy is 81%: the ability profile largely
    predicts occupation preference, but not fully (occupation choice is multi-factorial -- personality and interest
    matter too). One of Random Forest's most valuable outputs is the "variable importance ranking", telling which
    subject matters most for the choice. It is a powerful, easy-to-build machine-learning method for nonlinear,
    interacting relationships.

====================================================================================

#64  Support Vector Machine (SVM)
    file: 64_svm_university_passed.xlsx
  >> SCENARIO (narration):
    Imagine we want to predict whether a student will pass the university
    entrance exam from their academic indicators. For 250 students we recorded
    their GPA, math score, a practice trial_score, and weekly_study hours, with
    the outcome university_passed coded as 1 for passed and 0 for not passed. A
    Support Vector Machine suits this binary classification problem well because
    it finds the optimal boundary separating successful from unsuccessful
    students, and with a kernel it can capture nonlinear separations between the
    two groups even when the predictors overlap.
  >> VARIABLE SELECTION:
    - Target (binary): university_passed (1 = passed, 0 = not passed)
    - Predictors: GPA
    - Predictors: math
    - Predictors: trial_score
    - Predictors: weekly_study

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (university passed/not)
    Predicted from 4 variables (GPA, math, trial score, weekly study)

>> COMMENTARY (narration):
    We found the best boundary separating students who got into university from those who did not with SVM. SVM seeks
    the maximum-margin separating surface between two classes and, via the kernel trick, can model nonlinear
    boundaries too. Accuracy is 100% -- GPA, math, trial score and study separate the two groups perfectly (such high
    accuracy should be kept in mind regarding overfitting). SVM is especially powerful in small-to-medium,
    high-dimensional classification problems and is robust thanks to margin maximisation. It is a strong alternative
    to Random Forest.

====================================================================================

#65  Gradient Boosting
    file: 65_gb_school_dropout_risk.xlsx
  >> SCENARIO (narration):
    Suppose a school wants to flag students at risk of dropping out so it can
    intervene early. For 280 students we have GPA, absenteeism, a sound
    (noise/environment) indicator, and a parent_joint involvement measure, with
    the target school_dropout_risk classified as either low or medium. Gradient
    Boosting is well suited here because it builds an ensemble of trees
    sequentially, each correcting the previous one's errors, which yields strong
    predictive accuracy and a clear importance ranking of the risk factors
    driving dropout.
  >> VARIABLE SELECTION:
    - Target (class): school_dropout_risk (low / medium)
    - Predictors: GPA
    - Predictors: absenteeism
    - Predictors: sound
    - Predictors: parent_joint

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.964   (school dropout risk classification)
    GPA + absenteeism + socio-economic + parent togetherness

>> COMMENTARY (narration):
    We predicted school dropout risk from environmental/academic variables with Gradient Boosting. Boosting adds
    trees sequentially, each new tree correcting the previous ones' errors; accuracy is very high at 96.4%. This is
    critical for early-warning systems: the model can identify at-risk students with high accuracy. While Random
    Forest builds trees in parallel, Boosting builds them sequentially (chasing the error); it usually gives higher
    accuracy but is more prone to overfitting and needs careful tuning. It is a valuable tool for anticipating
    dropout and intervening.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_student_4kume.xlsx
  >> SCENARIO (narration):
    Let's say we want to discover natural groupings of students based on their
    academic behavior, without any predefined labels. For 240 students we have
    GPA, weekly study_hour, math score, and a trial_score. K-Means Clustering is
    the right tool because it partitions students into a chosen number of
    homogeneous clusters by minimizing within-group distance, helping us
    identify profiles such as high-achieving hard workers or low-engagement
    learners purely from the data.
  >> VARIABLE SELECTION:
    - Clustering variables: GPA
    - Clustering variables: study_hour
    - Clustering variables: math
    - Clustering variables: trial_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.458 (moderate cluster structure)
    Student typology from 4 variables (GPA, study, math, trial)

>> COMMENTARY (narration):
    We split students into four natural groups (a typology) by their academic features with K-Means. Silhouette =
    0.46 is moderate -- the clusters separate but not sharply, which is expected in educational data (achievement is
    a continuous gradient, sharp breaks are rare). The four groups likely represent distinct student profiles:
    high-achieving hard-workers, low-achieving low-effort, and intermediate types. K-Means finds natural groups in
    unlabeled data; it is practical for student profiling, designing targeted interventions and identifying risk groups.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5stil_10davranis.xlsx
  >> SCENARIO (narration):
    Imagine we want to see how a small group of students naturally group
    together based on their classroom behaviors, and to visualize the nested
    structure of those groups. For 40 students we have ten behavioral measures,
    behavior_01 through behavior_10, and a recorded learning style for later
    interpretation. Hierarchical Clustering is appropriate because it builds a
    dendrogram showing how students merge step by step, letting us decide the
    number of clusters afterward and inspect whether behavioral groupings align
    with learning styles.
  >> VARIABLE SELECTION:
    - Clustering variables: behavior_01
    - Clustering variables: behavior_02
    - Clustering variables: behavior_03
    - Clustering variables: behavior_04
    - Clustering variables: behavior_05
    - Clustering variables: behavior_06
    - Clustering variables: behavior_07
    - Clustering variables: behavior_08
    - Clustering variables: behavior_09
    - Clustering variables: behavior_10

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 5   Silhouette = 0.131 (weak/overlapping structure)
    5 styles from 10 behavior items

>> COMMENTARY (narration):
    We grouped student behaviors by similarity with hierarchical clustering -- the result is a dendrogram (tree).
    Splitting into five clusters gives a silhouette of only 0.13: the clusters overlap, there is no sharp separation.
    This is meaningful in education -- behavior styles are often a continuous spectrum, not sharp categories. A low
    silhouette means "natural grouping in the data is weak", which is informative (perhaps fewer than 5 styles
    exist). Hierarchical clustering is more explanatory than K-Means for discovering the nested structure of groups
    and the right number of clusters.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan_school_location.xlsx
  >> SCENARIO (narration):
    Suppose we have the geographic locations of 140 schools and we want to find
    where they cluster densely and which schools sit isolated as outliers. Each
    school has latitude (lat) and longitude (lon) coordinates. DBSCAN is the
    natural choice because it groups points by spatial density rather than a
    fixed number of clusters, automatically detecting dense pockets of schools
    while labeling sparsely located ones as noise, which is exactly what we need
    for mapping educational provision.
  >> VARIABLE SELECTION:
    - Coordinates: lat
    - Coordinates: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.878   (school locations)
    Schools clustered by geographic density

>> COMMENTARY (narration):
    We clustered schools' geographic locations by density with DBSCAN. Unlike K-Means: we do not give the number of
    clusters in advance, and it flags low-density isolated schools as "noise" (outliers). Four dense school groups
    were found (silhouette 0.88, very strong). This shows schools cluster into city centres/regions. DBSCAN is more
    appropriate than K-Means for spatial data with irregularly shaped clusters and outliers (school locations,
    student addresses); it is used to analyse the geographic distribution of educational services and access gaps.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca_academic_6ders.xlsx
  >> SCENARIO (narration):
    Let's say we measured students across six subjects and suspect these scores
    share a few underlying dimensions of ability. For 180 students we have math,
    science, language, reading_comprehension, logic, and art scores. Principal
    Component Analysis is well suited here because it reduces these six
    correlated subjects into a smaller set of uncorrelated components, revealing
    for instance a general academic factor and perhaps a verbal-versus-
    quantitative contrast, which simplifies interpretation and visualization.
  >> VARIABLE SELECTION:
    - Variables: math
    - Variables: science
    - Variables: language
    - Variables: reading_comprehension
    - Variables: logic
    - Variables: art

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 = 87.5%   (the first component alone explains most of the variance)
    6 subjects (math, science, language, reading, logic, art)

>> COMMENTARY (narration):
    We reduced six subject scores to a few summary axes with PCA. The first component alone explains 87.5% of the
    variance -- strikingly high. The meaning: subjects are very highly correlated, all measuring a single "general
    academic ability" dimension (g-factor-like). So a student good at math is generally good at science and language
    too. PCA is the basic tool for summarising multivariate data, reducing collinearity and revealing hidden
    dimensions (like general ability). One component being this dominant is the educational reflection of the
    "general intelligence/ability" debate.

====================================================================================

#70  t-SNE
    file: 70_tsne_student_5tip.xlsx
  >> SCENARIO (narration):
    Imagine we collected responses on twelve survey items from 125 students and
    we want to visualize whether students fall into distinct types in a two-
    dimensional map. We have item_01 through item_12, plus a known student_type
    label (A, B, C, D, E) we can color the plot by. t-SNE is ideal because it is
    a nonlinear dimension-reduction technique that preserves local neighborhood
    structure, so similar students appear close together, making it easy to see
    whether the five student types separate visually in the embedding.
  >> VARIABLE SELECTION:
    - Variables: item_01
    - Variables: item_02
    - Variables: item_03
    - Variables: item_04
    - Variables: item_05
    - Variables: item_06
    - Variables: item_07
    - Variables: item_08
    - Variables: item_09
    - Variables: item_10
    - Variables: item_11
    - Variables: item_12
    - Color/label (optional): student_type (A / B / C / D / E)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence (fit quality) = 0.519 -> good
    12-item profile (5 student types) reduced to 2D

>> COMMENTARY (narration):
    t-SNE is a nonlinear method that projects high-dimensional student profile data (12 items) into a two-dimensional
    plot. Unlike PCA, it focuses on preserving local neighbourhoods -- similar students are placed close on the map
    -- ideal for seeing hidden group structure. The KL divergence is 0.52, so the reduction quality is good; the five
    student types form visible clusters on the map. Important caveat: t-SNE axes and between-cluster distances are not
    interpretable, only the clustering pattern is. It is a powerful tool for visually exploring student types.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds_student_distance.xlsx
  >> SCENARIO (narration):
    Suppose we have 51 students described by their academic profile - GPA,
    weekly study time, and their math, science, and language scores. We want to
    map how similar or different students are to one another, so that students
    with comparable profiles sit close together on a two-dimensional plot.
    Multidimensional Scaling is the right tool here because it takes the
    pairwise distances between students across all five academic features and
    reconstructs a low-dimensional spatial map that preserves those distances as
    faithfully as possible, letting us visually spot clusters of academically
    similar learners.
  >> VARIABLE SELECTION:
    - Coordinates / Variables: GPA
    - Coordinates / Variables: study
    - Coordinates / Variables: math
    - Coordinates / Variables: science
    - Coordinates / Variables: language

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.054 (acceptable/good fit)
    Student similarity map (GPA, study, 3 subjects)

>> COMMENTARY (narration):
    MDS maps the similarity among students into two dimensions -- students with similar profiles placed near,
    dissimilar ones far. Stress = 0.054 is a "good" fit; this means the student-similarity structure fits two
    dimensions successfully (low stress = reliable map). Students close on the map have similar academic profiles.
    MDS is the classic way to visualise multivariate similarity/distance data; the stress value honestly tells how
    much we can trust the map. It is used to see student grouping and profile similarity.

====================================================================================

#72  UMAP
    file: 72_umap_behavior_5tip.xlsx
  >> SCENARIO (narration):
    Here we have 200 students rated on 25 behavioral indicators, and each
    student also belongs to one of five latent behavior profiles, labeled S1
    through S5 in lower_type. We want to see whether these 25 behaviors
    naturally separate the students into the five known profile groups when
    projected onto a 2D plane. UMAP is ideal for this because it is a nonlinear
    dimensionality reduction technique that preserves local neighborhood
    structure, so behaviorally similar students stay together, and we can color
    the embedding by lower_type to check whether the profiles form distinct
    clusters.
  >> VARIABLE SELECTION:
    - Features (high-dim input): behavior_01 ... behavior_25 (all 25 behavior items)
    - Color / Label (grouping): lower_type (S1, S2, S3, S4, S5)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    UMAP dimension reduction (25 behavior items -> 2D, 5 types)
    n_neighbors and min_dist balance local/global structure

>> COMMENTARY (narration):
    UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure
    better and is faster. We projected student profiles of 25 behavior items into a two-dimensional map; similar
    behavior profiles cluster. The n_neighbors parameter tunes the local-global balance, and min_dist the tightness
    of clusters. UMAP is increasingly preferred for discovering hidden group structure in high-dimensional
    behavior/survey data. It is used for visualisation; the clustering pattern is interpreted, not the axis values.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_school_survival_21.xlsx
  >> SCENARIO (narration):
    We administered a 21-item school adaptation and survival scale to 200
    students, with every item on the same response metric. Before we use the
    total scale score in any further analysis, we need to confirm that all 21
    items are measuring the same underlying construct consistently. Cronbach's
    Alpha is the appropriate reliability index here because it estimates the
    internal consistency of a set of items intended to form a single scale, and
    the item-total statistics also tell us whether dropping any item would
    improve reliability.
  >> VARIABLE SELECTION:
    - Items (scale): item_01
    - Items (scale): item_02
    - Items (scale): item_03
    - Items (scale): item_04
    - Items (scale): item_05
    - Items (scale): item_06
    - Items (scale): item_07
    - Items (scale): item_08
    - Items (scale): item_09
    - Items (scale): item_10
    - Items (scale): item_11
    - Items (scale): item_12
    - Items (scale): item_13
    - Items (scale): item_14
    - Items (scale): item_15
    - Items (scale): item_16
    - Items (scale): item_17
    - Items (scale): item_18
    - Items (scale): item_19
    - Items (scale): item_20
    - Items (scale): item_21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.974 (excellent internal consistency)   21 items
    School life/belonging scale

>> COMMENTARY (narration):
    We tested the internal consistency of a 21-item school life/belonging scale with Cronbach's alpha. Alpha = 0.97
    is excellent: the 21 items consistently measure the same underlying concept, and students answer coherently
    across items. Showing a scale's reliability before analysing it is essential -- low alpha signals the items do
    not measure the same thing and the total score may be meaningless. Above 0.70 is acceptable, above 0.90
    excellent. It is the first quality step of any educational study developing a scale. (A very high alpha can also
    hint at item redundancy.)

====================================================================================

#74  Likert Analysis
    file: 74_likert_student_3boyut.xlsx
  >> SCENARIO (narration):
    We collected responses from 220 students on a Likert questionnaire built
    around three attitude dimensions - five motivation items, five self-efficacy
    items, and five anxiety items. We want to summarize how students responded
    across each dimension, looking at the distribution of agreement levels and
    the mean response per item. Likert Analysis is the right approach because it
    is designed specifically for ordered rating-scale items, producing frequency
    breakdowns, diverging response patterns, and per-item summaries grouped by
    their underlying construct.
  >> VARIABLE SELECTION:
    - Items (motivation): motivation_1
    - Items (motivation): motivation_2
    - Items (motivation): motivation_3
    - Items (motivation): motivation_4
    - Items (motivation): motivation_5
    - Items (self-efficacy): self_efficacy_1
    - Items (self-efficacy): self_efficacy_2
    - Items (self-efficacy): self_efficacy_3
    - Items (self-efficacy): self_efficacy_4
    - Items (self-efficacy): self_efficacy_5
    - Items (anxiety): anxiety_1
    - Items (anxiety): anxiety_2
    - Items (anxiety): anxiety_3
    - Items (anxiety): anxiety_4
    - Items (anxiety): anxiety_5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.793 (acceptable)   15 items (3 dimensions: motivation/self-efficacy/anxiety)

>> COMMENTARY (narration):
    We examined the overall reliability and item statistics of a 15-item, three-dimensional (motivation,
    self-efficacy, anxiety) Likert scale. Cronbach's alpha = 0.79 is "acceptable" -- the scale is generally
    consistent, but because three different dimensions are combined, a single alpha can come out a bit low
    (sub-dimensions should be analysed separately). Likert analysis gives item means, distributions and item-total
    correlations to show whether the scale is sound. It is used for the basic quality control of attitude/motivation
    scales in education.

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18_item_3faktor.xlsx
  >> SCENARIO (narration):
    We have responses from 250 teachers on an 18-item questionnaire, q01 through
    q18, and we suspect the items tap into a smaller number of underlying
    attitude factors. We don't yet have a confirmed theoretical structure, so we
    want to let the data reveal how the items group together. Exploratory Factor
    Analysis is the suitable method here because it uncovers latent factors from
    the pattern of inter-item correlations, and the factor loadings will show us
    which items cluster onto which of the expected three dimensions.
  >> VARIABLE SELECTION:
    - Items (variables to factor): q01
    - Items (variables to factor): q02
    - Items (variables to factor): q03
    - Items (variables to factor): q04
    - Items (variables to factor): q05
    - Items (variables to factor): q06
    - Items (variables to factor): q07
    - Items (variables to factor): q08
    - Items (variables to factor): q09
    - Items (variables to factor): q10
    - Items (variables to factor): q11
    - Items (variables to factor): q12
    - Items (variables to factor): q13
    - Items (variables to factor): q14
    - Items (variables to factor): q15
    - Items (variables to factor): q16
    - Items (variables to factor): q17
    - Items (variables to factor): q18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO (overall) = 0.908 (Excellent)   -> factor analysis very appropriate
    18 items, expected 3 factors

>> COMMENTARY (narration):
    We explored how many hidden dimensions (factors) underlie 18 survey items with Exploratory Factor Analysis. KMO =
    0.91 is "excellent": the data are very suitable for factor analysis, with high shared variance among items. The
    items group into three factors as expected. EFA reduces many items to a few meaningful dimensions, simplifying
    the scale and revealing the underlying structures. It is a core tool of scale development and psychometric/
    educational research; KMO and Bartlett's test are reported as the pre-check of suitability.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3ogretmen_composition.xlsx
  >> SCENARIO (narration):
    Three teachers independently scored the same 50 student compositions,
    recorded as teacher_1, teacher_2, and teacher_3. Before we trust these
    ratings, we need to know how consistent the teachers are with one another.
    Intraclass Correlation is the right statistic because it quantifies the
    degree of agreement among multiple raters scoring the same subjects on a
    continuous scale, telling us what proportion of the score variance reflects
    true differences between compositions rather than rater disagreement.
  >> VARIABLE SELECTION:
    - Raters / Measurements: teacher_1
    - Raters / Measurements: teacher_2
    - Raters / Measurements: teacher_3

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.894   (Excellent)   3 teachers   95% CI ~ (0.84 , 0.93)

>> COMMENTARY (narration):
    We assessed how consistent three teachers are in scoring the same compositions/assignments with ICC. ICC = 0.89
    is "excellent": inter-teacher scoring agreement is very high -- whichever teacher scores, the result is similar.
    This is critical in educational assessment: if scoring varied from teacher to teacher, student grades would be
    unfair. ICC is the standard for measuring inter-rater reliability on continuous measurements (scores, ratings)
    (Cohen's Kappa is for categorical data, ICC for continuous). It is essential for rubric validity and grading
    consistency.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_self_academic_social.xlsx
  >> SCENARIO (narration):
    We measured 280 students on twelve indicators of self-concept, four for the
    academic self, four for the general self, and four for the social self.
    Unlike an exploratory analysis, here we already have a clear theoretical
    model: three correlated latent factors with specified item loadings.
    Confirmatory Factor Analysis is the appropriate method because it lets us
    test whether this hypothesized three-factor structure fits the observed
    data, reporting fit indices and factor loadings that confirm or challenge
    our measurement model.
  >> VARIABLE SELECTION:
    - Factor 1 (self) items: self_1, self_2, self_3, self_4
    - Factor 2 (academic) items: academic_1, academic_2, academic_3, academic_4
    - Factor 3 (social) items: social_1, social_2, social_3, social_4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 1.00   RMSEA = 0.000   (excellent fit)
    3 factors: self + academic + social

>> COMMENTARY (narration):
    While EFA explores hidden structure, CFA tests a pre-specified theory: here we tested whether three distinct
    self/competence dimensions -- "self", "academic" and "social" -- fit the data. The fit is excellent (CFI = 1.00,
    RMSEA = 0.000): the items really load onto these three dimensions, the hypothesised scale structure is fully
    confirmed. CFA is part of structural equation modelling (SEM) and is the strongest way to prove scale validity.
    It is the standard method for demonstrating the validity of multidimensional educational/psychological constructs
    like self-concept and self-efficacy.

====================================================================================

#78  Survey Means
    file: 78_survey_means_PISA.xlsx
  >> SCENARIO (narration):
    This is a national PISA reading survey of 350 students drawn from four
    regions - Marmara, the Southeast, Inner Anatolia, and the East - and the
    sample was collected with unequal probabilities, so each student carries a
    sampling weight. We want to estimate the average PISA reading score for the
    population, both overall and broken down by region. Survey Means is the
    correct procedure because it applies the design weights when computing the
    mean and its standard error, giving us unbiased population estimates rather
    than naive unweighted averages.
  >> VARIABLE SELECTION:
    - Analysis variable (mean): PISA_reading
    - Weight: weight
    - By / Domain (subgroup): region (Marmara, G.east, inner anatolia, east)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PISA reading weighted mean = 464.33   SE = 3.99   95% CI (456.48 , 472.18)   CV 0.86%
    Stratified + weighted design (Taylor SE)

>> COMMENTARY (narration):
    Large-scale exams like PISA select students not at random but with a stratified, weighted design. A simple mean
    would be wrong for such data; complex-sample analysis gives the correct mean and standard error (Taylor
    linearization) by accounting for the design weights and stratification. The PISA reading weighted mean is 464.3,
    with a very narrow confidence interval (CV 0.86%). National/international educational assessments (PISA, TIMSS,
    national monitoring) are always designed this way; survey methods are essential for unbiased, design-consistent
    estimation.

====================================================================================

#79  Survey Frequencies
    file: 79_survey_freq_university_intention.xlsx
  >> SCENARIO (narration):
    We surveyed 400 students about their intention to attend university, with
    answers ranging across none, maybe, probable, and certain, and students came
    from urban and rural regions under a weighted sampling design. We want to
    estimate the population proportion of students in each intention category.
    Survey Frequencies is the right tool because it produces weighted frequency
    and percentage estimates for categorical variables, properly accounting for
    the sampling weights so the percentages reflect the target population.
  >> VARIABLE SELECTION:
    - Categorical variable (frequencies): university_intention (none, maybe, probable, certain)
    - Weight: weight
    - By / Domain (subgroup): region (urban, rural)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    "certain" proportion = 0.050 (SE 0.011)   |   "maybe" proportion = 0.382 (SE 0.025)
    University intention -- stratified + weighted design

>> COMMENTARY (narration):
    We estimated the proportion of university-intention categories among students, accounting for the complex sample
    design. Those saying "maybe" are 38%, "certain" 5% (with design-based standard errors). A simple percentage would
    give the wrong standard error by ignoring stratification and weights; survey frequency analysis produces a
    confidence interval reflecting the true uncertainty. It is the correct way to estimate category proportions in a
    population (university intention, occupation preference, attitude distribution) in a design-consistent manner.

====================================================================================

#80  Survey Totals
    file: 80_survey_total_teacher.xlsx
  >> SCENARIO (narration):
    We have data on 280 schools, each reporting its teacher_count, drawn under a
    stratified design with small, medium, and large size strata and accompanying
    sampling weights. We want to estimate the total number of teachers across
    the whole population of schools, and within each size stratum. Survey Totals
    is the appropriate procedure because it produces weighted population total
    estimates with correct design-based standard errors, which is exactly what
    we need when scaling a sampled count up to the full population.
  >> VARIABLE SELECTION:
    - Analysis variable (total): teacher_count
    - Weight: weight
    - By / Stratum (subgroup): size_stratum (S-small, M-medium, L-large)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total teacher count = 69,470   SE = 1,207   95% CI (67,093 , 71,847)
    Size-stratified design (Taylor SE)

>> COMMENTARY (narration):
    We estimated the total number of teachers in a region using the sampling design and weights: about 69,470 (95% CI
    67,093 - 71,847). Scaling from the sampled schools up to the whole population requires correct weighting -- each
    sampled school carries weight proportional to the population fraction it represents. Survey total analysis does
    exactly this and expresses uncertainty with the standard error. It is the official-statistics method used to
    estimate educational population totals (total teachers, total students, total schools) from a sample.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg_PISA.xlsx
  >> SCENARIO (narration):
    Suppose we run a national survey of students and we want to model how their
    PISA scores relate to background characteristics. We have 320 students with
    their age, the noise or sound level in their study environment, and the
    region they live in, either east or Marmara. Because the data come from a
    complex sampling design, each student carries a survey weight, so an
    ordinary regression would give biased estimates of the population
    relationship. We use survey-weighted linear regression to predict PISA_score
    from age, sound, and region while properly accounting for the design
    weights, giving us population-representative coefficients.
  >> VARIABLE SELECTION:
    - Dependent variable: PISA_score
    - Predictors: age, sound, region (east/Marmara)
    - Weight: weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.167   age coefficient = 1.26   p = 0.69 (non-significant)
    PISA score ~ age + socio-economic level   (stratified + weighted)

>> COMMENTARY (narration):
    We modelled PISA score from age and socio-economic level, accounting for the complex sample design (strata,
    weights). The age effect is non-significant (p = 0.69) -- in this age range, age does not predict PISA score. The
    model explains 17% of the variance. Design-based standard errors are more accurate than the (often too
    optimistic) errors of simple OLS. When estimating relationships from national/international assessment data (where
    the sample is not random), survey regression is necessary for valid inference; otherwise p-values would mislead.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic_university.xlsx
  >> SCENARIO (narration):
    Here we want to understand which students pass the university entrance exam,
    again from a weighted national sample. For each of 320 students we record
    their GPA, the sound level of their study environment, the region they come
    from, either Marmara or east, and a binary indicator of whether they passed,
    university_passed. Since the outcome is yes or no and the sample is drawn
    with unequal selection probabilities, we use survey-weighted logistic
    regression. This lets us estimate how GPA and the other predictors change
    the population-level odds of passing the exam while respecting the sampling
    weights.
  >> VARIABLE SELECTION:
    - Dependent variable (binary): university_passed
    - Predictors: GPA, sound, region (Marmara/east)
    - Weight: weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    GPA: OR = 5.68   p < .001 ***   95% CI (3.02 , 10.67)
    University admission ~ GPA + socio-economic   (stratified + weighted)

>> COMMENTARY (narration):
    Predicting a student's probability of university admission from GPA, we handled both the binary outcome
    (logistic) and the complex sample design together. The odds ratio is 5.68 (p < .001): a one-unit rise in GPA
    multiplies the odds of admission 5.7-fold -- a very strong predictor. Because the design weights and
    stratification are accounted for, the OR and confidence interval are design-consistent and valid. Survey logistic
    regression is the correct way to model binary outcomes (admitted/not, graduated/not) from complex educational
    surveys.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_study_achievement.xlsx
  >> SCENARIO (narration):
    Imagine we suspect that the relationship between how many hours students
    study and how well they score on a test is not a straight line. Maybe gains
    accelerate at first and then flatten out as study time grows. With 200
    students and their weekly study_hour and test_score, a plain linear
    regression would miss that curvature. We use a Generalized Additive Model,
    which fits a smooth, flexible spline of study_hour onto test_score and lets
    the data reveal the true shape of the relationship instead of forcing it to
    be linear.
  >> VARIABLE SELECTION:
    - Dependent variable: test_score
    - Smooth predictor: study_hour

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (explained) = 0.615
    Test score ~ s(study hours)  (smooth function)

>> COMMENTARY (narration):
    GAM is a flexible generalisation of linear regression: it models a variable's effect with "smooth curves"
    (splines) rather than a straight line. Study's effect on test score is often nonlinear -- it may show diminishing
    returns (study saturation) beyond a point. GAM learns this curved pattern from the data without assuming its
    shape in advance; the model explains 62% of the variance. It preserves interpretability while adding flexibility:
    we can read from the curve where study helps most. For nonlinear but interpretable relationships in education
    (study-achievement, age-development), it is an ideal method.

====================================================================================

#84  Discriminant Analysis (LDA/QDA)
    file: 84_diskriminant_3sinif_5feat.xlsx
  >> SCENARIO (narration):
    Suppose we want to classify students into three achievement levels, low,
    medium, and high, based on their academic profile. For 105 students we have
    their math and science scores, GPA, absenteeism days, and weekly study
    hours. We want both to know which combination of variables best separates
    the three groups and to build a rule that assigns a new student to a level.
    We use Discriminant Analysis, which finds the linear combinations of these
    continuous predictors that maximally distinguish the achievement classes and
    classifies cases accordingly.
  >> VARIABLE SELECTION:
    - Grouping variable (target): achievement_class (low/medium/high)
    - Predictors: math, science, GPA, absenteeism, study_hour

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (3 achievement classes)
    Separation from 5 variables (math, science, GPA, absenteeism, study)

>> COMMENTARY (narration):
    Discriminant analysis finds the linear combinations of variables that best separate the groups (three achievement
    classes) -- a bit like a classification-focused PCA. Accuracy is 100%: five variables separate the three classes
    perfectly. Discriminant analysis both classifies and answers "which variables matter most for the separation"
    (discriminant function loadings). It resembles logistic regression but is the classic choice for multiple groups
    when variables are normally distributed. It is used in achievement-level classification and for assigning new
    students to existing groups.

====================================================================================

#85  Conditional Logit (Choice Model)
    file: 85_conditional_logit_university.xlsx
  >> SCENARIO (narration):
    Here we study how students choose among universities. In a discrete choice
    experiment, each of our students faces a set of four alternatives,
    U1-public-near, U2-public-far, U3-private-near, and U4-private-far, and
    picks one. Every alternative is described by its annual_fee and its
    distance_km, and chosen marks which one was actually selected. Because the
    choice is among competing options whose attributes vary, an ordinary
    logistic model will not do. We use a Conditional Logit choice model, which
    estimates how fee and distance drive the probability that an alternative is
    chosen, with choice_id grouping the alternatives that belong to the same
    decision.
  >> VARIABLE SELECTION:
    - Choice indicator (event chosen): chosen
    - Choice set / group id: choice_id
    - Alternative-specific attributes: annual_fee, distance_km
    - Alternative label: university (U1-public-near/U2-public-far/U3-private-near/U4-private-far)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (McFadden) = 0.965  (very high)
    University choice ~ annual fee + distance (km)

>> COMMENTARY (narration):
    We examined which university students choose with a conditional logit model -- each student selects one from a
    choice set. McFadden pseudo R2 = 0.97 is very high (0.2-0.4 is already strong for choice models): fee and distance
    explain the choice almost completely. The model shows students strongly prefer nearby, lower-fee universities.
    This method underlies "discrete choice" analysis in transport, marketing and education economics (school/
    university choice, demand modelling); it is powerful for modelling access and cost effects in education policy.

====================================================================================

#86  Kaplan-Meier Survival Analysis
    file: 86_km_diploma_duration.xlsx
  >> SCENARIO (narration):
    Suppose we want to describe how long it takes students to graduate and earn
    their diploma. For 200 students we record duration_term, the number of terms
    observed, and graduate, which is 1 if they actually graduated and 0 if they
    were still enrolled or left before graduating, so their time is censored. We
    use Kaplan-Meier survival analysis to estimate the graduation curve over
    time, and by splitting on department, either education, social, eng, or
    health, we can compare how quickly students in different faculties reach
    graduation.
  >> VARIABLE SELECTION:
    - Time variable: duration_term
    - Event indicator (1=graduated): graduate
    - Group/strata: department (education/social/eng/health)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival (time to graduation) = 8.2 terms   (grouped by department)
    Log-rank test compares groups

>> COMMENTARY (narration):
    We analysed students' time to graduation -- by department -- with Kaplan-Meier. Median time to graduation is 8.2
    terms (about 4 years). Kaplan-Meier analyses "time-to-event" data -- correctly using not-yet-graduated (censored)
    students too -- and tests department differences in duration with the log-rank test. In education it is the basic
    visual and test for "time-event" data such as time to graduation, time to dropout, diploma attainment; it is at
    the centre of student-progress analysis.

====================================================================================

#87  Cox Proportional Hazards Regression
    file: 87_cox_dropout_hazard.xlsx
  >> SCENARIO (narration):
    Here we want to know which factors raise or lower a student's risk of
    dropping out over time. For 250 students we have duration_term as the
    follow-up time and dropout as the event, along with covariates: age, GPA,
    whether they hold a scholarship, and the level of family_support. We use Cox
    proportional hazards regression because it models the instantaneous risk of
    dropout as a function of these predictors without assuming a particular
    baseline shape, telling us, for example, how much a scholarship reduces the
    hazard of leaving school.
  >> VARIABLE SELECTION:
    - Time variable: duration_term
    - Event indicator (1=dropout): dropout
    - Covariates: age, GPA, scholarship, family_support

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.638   GPA: HR = 0.581   95% CI (0.394 , 0.858)   p = 0.006 **
    Dropout ~ age + GPA + scholarship + family support

>> COMMENTARY (narration):
    We modelled a student's dropout risk from academic/demographic variables with Cox regression. GPA is significant
    and protective (HR = 0.58 < 1, p = 0.006): a one-unit rise in GPA roughly halves the dropout risk -- high
    achievement strongly encourages staying in school. Cox regression is the gold-standard survival model relating
    "time-to-event" to several predictors; it gives hazard ratios (HR). Anticipating dropout and identifying risk
    factors -- for early intervention -- is very valuable in education. HR < 1 is protective, HR > 1 risk-increasing.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT) Model
    file: 88_aft_weibull_mezuniyet.xlsx
  >> SCENARIO (narration):
    Suppose that instead of modeling the risk of an event, we want to model the
    actual time until students graduate and ask how a covariate stretches or
    shortens that time. For 200 students we observe duration_term, the event
    indicator graduate, and whether the student had a scholarship. We use a
    parametric Accelerated Failure Time model with a Weibull distribution, which
    directly estimates how scholarship accelerates or decelerates the time to
    graduation, giving an interpretable time-ratio rather than a hazard ratio.
  >> VARIABLE SELECTION:
    - Time variable: duration_term
    - Event indicator (1=graduated): graduate
    - Covariate: scholarship

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution = Weibull   Median survival = 7.35 terms   AIC = 994.12
    Time to graduation ~ scholarship

>> COMMENTARY (narration):
    We analysed time to graduation with a parametric Weibull survival model. While Cox is semi-parametric, Weibull
    models the whole survival curve with a specific mathematical distribution -- if the data fit it, it gives more
    efficient estimates and allows extrapolation. Median time to graduation is 7.35 terms. The AFT (accelerated
    failure time) interpretation is intuitive: how many times a variable (like scholarship) lengthens/shortens the
    time to graduate. When the time distribution is known, parametric models give more powerful and predictable
    results than Cox.

====================================================================================

#89  Competing Risks
    file: 89_competing_risks_3neden.xlsx
  >> SCENARIO (narration):
    Imagine students can leave the program for several different reasons, and we
    do not want to lump them together. For 220 students we have duration_term,
    an event indicator, and separation_cause coding the distinct competing
    reasons for leaving, with age_covariate as a predictor. Because one type of
    exit prevents the others from happening, treating non-target exits as simple
    censoring would overstate risk. We use a Competing Risks analysis to
    estimate the cumulative incidence of each separation cause separately while
    accounting for the others.
  >> VARIABLE SELECTION:
    - Time variable: duration_term
    - Event/status indicator: event
    - Competing event type: separation_cause
    - Covariate: age_covariate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 130   |   3 different separation causes (event types)
    Separate cumulative incidence (CIF) and risk for each event type

>> COMMENTARY (narration):
    A student can leave school for more than one reason: academic failure, economic, or transfer -- and these are
    mutually exclusive (once one happens the others cannot). Classic survival handles this incorrectly; competing-
    risks analysis produces a separate cumulative incidence for each separation cause. In the data 130 students are
    still enrolled (censored), the rest left for different reasons. This framework correctly answers questions like
    "which factor increases the risk of which separation cause". When there are different dropout pathways in
    education, competing-risks analysis is the correct and informative method.

====================================================================================

#90  Time-Dependent Cox Model
    file: 90_tvcox_GPA_dropout.xlsx
  >> SCENARIO (narration):
    Suppose a student's GPA changes from term to term, and we believe their
    current GPA, not just their starting GPA, drives dropout risk. Our data are
    in counting-process form with 464 records, each giving a start and end term
    interval, the GPA that held during that interval, and an event flag for
    whether dropout occurred at its end, with student_id linking a student's
    successive intervals. We use a Time-Dependent Cox model, which lets GPA vary
    over follow-up so the hazard at any moment reflects the student's GPA at
    that moment.
  >> VARIABLE SELECTION:
    - Start time: start
    - Stop time: end
    - Event indicator (1=dropout): event
    - Time-varying covariate: GPA
    - Subject id: student_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    GPA: HR = 1.158   95% CI (0.609 , 2.206)   p = 0.654 (non-significant)
    Time-varying GPA covariate

>> COMMENTARY (narration):
    A student's GPA changes from term to term; standard Cox assumes it constant, time-dependent Cox uses the current
    value in each time interval (start-stop format). In this model the effect of current GPA on dropout risk is
    non-significant (HR = 1.16, p = 0.65) -- in this data the momentary GPA change does not predict dropout timing. A
    non-significant result is information too. Time-dependent Cox is the correct method when "the current,
    continuously monitored condition matters, not just the baseline"; it is the standard for modelling time-varying
    factors (GPA, absenteeism, support) in longitudinal educational monitoring.

====================================================================================

#91  Survey-Weighted Cox Regression
    file: 91_survey_phreg_dropout.xlsx
  >> SCENARIO (narration):
    Suppose we are studying why some students drop out of university and we want
    our hazard model to reflect the national student population, not just our
    raw sample. We followed 300 students over time, recording how many terms
    they stayed enrolled in duration_term and whether they eventually dropped
    out, coded by dropout. We also know each student's region (A, B, or C),
    their age, and the school they attended. Because the data came from a
    complex sampling design, every student carries a survey weight, so an
    ordinary Cox model would give biased estimates. We use Survey-Weighted Cox
    Regression so the proportional-hazards estimates for dropout properly
    account for these design weights and generalize to the wider population.
  >> VARIABLE SELECTION:
    - Time (duration): duration_term
    - Event: dropout (1=dropped out)
    - Weight: weight
    - Predictors (covariates): age, region (A/B/C)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.53   Weighted/clustered survival (age covariate, school cluster)
    Dropout ~ age

>> COMMENTARY (narration):
    We combined survival analysis with a complex sample design: students come from a population sampled with weights
    and clustered by schools. Survey-PHREG runs the Cox model with design weights and cluster-robust standard errors,
    so the hazard ratios and confidence intervals generalise to the population. Concordance 0.53 means age alone is a
    weak discriminator in this model. When national education surveys have "time-to-event" data, ignoring the design
    biases the estimates; survey-PHREG provides valid inference.

====================================================================================

#92  Interval-Censored Survival Analysis
    file: 92_interval_censored_sinav.xlsx
  >> SCENARIO (narration):
    Imagine we want to know how long it takes students to pass a key qualifying
    exam, but we never observe the exact term they pass. Instead, for each of
    our 60 students we only know an interval: the last term we checked and they
    had not yet passed (left_censor_term) and the term by which we confirmed
    they had passed (survival_censor_term). This kind of inexact, interval-
    bounded timing is common when assessments are only given periodically. We
    also recorded each student's department: engineering, social, or health. We
    use Interval-Censored Survival Analysis because the event time is only known
    to fall within an interval, and standard methods that assume exact times
    would be inappropriate.
  >> VARIABLE SELECTION:
    - Left interval (lower bound): left_censor_term
    - Right interval (upper bound): survival_censor_term
    - Group / covariate: department (eng/social/health)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 47   Median survival = 8.0 terms   (Turnbull algorithm)
    Event time known within [lower, upper] interval

>> COMMENTARY (narration):
    In educational monitoring we rarely know the exact time of an event: we check students at term ends, find them
    successful one term and over a threshold at the next check -- the event happened somewhere between. Interval-
    censored survival (Turnbull algorithm) handles exactly this uncertainty, avoiding fixing the event to an
    arbitrary date. Median survival is 8 terms. By the nature of periodic assessment (term exams, annual checks),
    event times are always within an interval; forcing this data to a single point biases the result. The
    interval-censored method matches the reality of monitoring data.

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty_cox_school.xlsx
  >> SCENARIO (narration):
    Suppose we are again modeling time to dropout, but this time our 120
    students are nested within different schools, and we suspect that students
    from the same school share unobserved risk factors, such as school climate
    or local support. Ignoring this clustering would understate the uncertainty
    in our estimates. For each student we have the number of terms enrolled in
    duration_term, the dropout indicator, their age, and a clinical_group label
    (K1 to K4), plus the school_id that defines the cluster. We use a Frailty
    Cox Model, which adds a random school-level frailty term to the Cox model so
    we can estimate covariate effects while accounting for shared, school-
    specific risk.
  >> VARIABLE SELECTION:
    - Time (duration): duration_term
    - Event: dropout
    - Frailty / cluster: school_id
    - Predictors (covariates): age, clinical_group (K1/K2/K3/K4)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.564   School-clustered shared frailty (random effect)
    Dropout ~ age  +  (school frailty)

>> COMMENTARY (narration):
    Students are clustered within schools; students in the same school share unmeasured common factors (school
    climate, teacher quality, resources) and carry similar dropout risk. Frailty Cox adds a shared "frailty" (random
    effect) for each school to model this clustering -- the mixed-model version of survival analysis. Concordance
    0.56. Accounting for school-level hidden differences gives both correct standard errors and information on "how
    much heterogeneity exists between schools". It is the correct method for clustered educational survival data
    (schools, classes).

====================================================================================

#94  Time Series Description
    file: 94_ts_monthly_kutuphane.xlsx
  >> SCENARIO (narration):
    Let's say our university library wants a clear summary of how monthly visits
    have behaved over the past five years. We have 60 consecutive monthly
    observations, with each month's date and the corresponding monthly_visit
    count. Before fitting any forecasting model, it is good practice to first
    describe the series: its overall level, range, whether it trends up or down,
    and whether it shows obvious seasonal swings around exam periods. We use
    Time Series Description to compute these summary characteristics and
    visualize the monthly_visit series indexed by date.
  >> VARIABLE SELECTION:
    - Time index (date): date
    - Series (value): monthly_visit

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.996 -> NOT stationary   Trend = increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly library-visit series. The ADF test does not find it stationary (p = 0.996) -- the mean
    changes over time, there is a trend (increasing) and seasonality is present (fluctuation tied to academic terms).
    This pre-diagnosis is critical: most time-series models (like ARIMA) require stationarity, so differencing/
    transformation may be needed first. Time-series analysis reveals the structure of long monitoring data -- trend,
    seasonality, stationarity -- and tells which modelling steps are required. It is the foundation for analysing
    educational usage/demand data (library, applications, absenteeism).

====================================================================================

#95  STL Decomposition
    file: 95_stl_motivation_daily.xlsx
  >> SCENARIO (narration):
    Suppose we collected the average daily student motivation score over three
    full school years, giving us 1095 daily observations of mean_motivation
    indexed by date. We can clearly see motivation rises and falls within each
    year, but we want to separate the underlying long-term trend from the
    recurring seasonal pattern and the day-to-day noise. We use STL
    Decomposition, which splits the daily motivation series into trend,
    seasonal, and remainder components, making it easy to see, for example,
    whether the trend is genuinely improving once we strip away predictable
    seasonal cycles.
  >> VARIABLE SELECTION:
    - Time index (date): date
    - Series (value): mean_motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Period (seasonal length) = 7 (daily -> weekly cycle)
    Series decomposed into trend + seasonal + residual components

>> COMMENTARY (narration):
    STL (Seasonal-Trend decomposition using Loess) splits a daily motivation series into three components: long-term
    trend, a recurring cycle (7-day = weekly) and the remaining residual. The weekly period is meaningful --
    motivation fluctuates regularly across the week (e.g. low at the start of the week, recovering later). This
    decomposition clarifies "is motivation generally rising, or just fluctuating weekly". STL is the most intuitive
    way to interpret seasonal/cyclical educational series, separating trend from cycle so each can be assessed
    separately.

====================================================================================

#96  ARIMA Forecasting
    file: 96_arima_monthly_graduate.xlsx
  >> SCENARIO (narration):
    Imagine the registrar wants to forecast how many students will graduate in
    the coming months. We have 120 months of historical data, with each month's
    date and its monthly_graduate count. The series shows both a trend and
    seasonal structure tied to the academic calendar. We use ARIMA Forecasting
    because it models the autocorrelation and seasonality in the graduate counts
    and produces forward forecasts with confidence intervals, helping the
    administration plan ceremonies and resources.
  >> VARIABLE SELECTION:
    - Time index (date): date
    - Series (value): monthly_graduate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model fitted   AIC = 2520.94
    Monthly graduate count -- future forecast

>> COMMENTARY (narration):
    We fitted an ARIMA model to a monthly graduate-count series and forecast the future. ARIMA combines the series'
    own past values (AR), past errors (MA) and differencing (I, to make it stationary). With AIC = 2521 the best
    model was selected. Forecasts rest on the observed trend and autocorrelation structure and come with an
    uncertainty band. Time-series forecasting matters for educational planning: anticipating future graduate/
    application counts enables quota, staffing and resource planning. ARIMA is the classic forecasting method for
    educational series with seasonality and trend.

====================================================================================

#97  Exponential Smoothing (Holt-Winters)
    file: 97_ets_monthly_application.xlsx
  >> SCENARIO (narration):
    Suppose the admissions office wants to anticipate the number of monthly
    applications for the next year. We have 84 months of application history,
    with each month's date and its monthly_application total, and the data
    clearly shows a rising trend plus a strong yearly seasonal peak around
    application deadlines. We use Exponential Smoothing with the Holt-Winters
    method because it explicitly models level, trend, and seasonality, making it
    well suited for forecasting this kind of seasonal application series.
  >> VARIABLE SELECTION:
    - Time index (date): date
    - Series (value): monthly_application

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method = Holt-Winters (additive seasonal, period 12)   AIC = 969.41
    Monthly application count forecast

>> COMMENTARY (narration):
    Holt-Winters exponential smoothing estimates the level + trend + seasonality components in a weighted way (giving
    more weight to the recent past). It captured the seasonal pattern of monthly applications (AIC = 969).
    Applications typically show strong seasonality (registration periods); Holt-Winters is often very successful for
    such series and is more intuitive than ARIMA. In educational demand forecasting (applications, enrolment,
    attendance) it is an easy-to-build, interpretable alternative, especially preferred for regularly seasonally
    fluctuating variables.

====================================================================================

#98  Mann-Kendall Trend and Sen's Slope
    file: 98_mann_kendall_reading_trend.xlsx
  >> SCENARIO (narration):
    Let's say we want to know whether national reading achievement has genuinely
    improved over the last four decades. We have 40 yearly observations, each
    with a year and the national_reading_score for that year. Rather than assume
    a straight-line trend, we want a robust, distribution-free test of whether
    there is a monotonic upward or downward tendency, plus an estimate of the
    rate of change. We use the Mann-Kendall trend test together with Sen's
    slope, because they detect and quantify a monotonic trend in the reading
    scores without requiring the data to be normally distributed.
  >> VARIABLE SELECTION:
    - Time (ordering): year
    - Series (value): national_reading_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   Trend = INCREASING
    National reading score long-term trend

>> COMMENTARY (narration):
    We tested whether national reading scores show a significant long-term trend with the Mann-Kendall test -- a
    nonparametric, outlier- and distribution-robust trend test. The result is highly significant (p < .001) and the
    direction is increasing: reading achievement is rising consistently over the years. This may reflect the positive
    cumulative effect of educational policies/interventions. Mann-Kendall + Sen's slope is the gold standard for
    long-term trend detection -- from climate to economy, from education to health; Sen's slope gives the trend
    magnitude without being affected by outliers.

====================================================================================

#99  Anomaly Detection
    file: 99_anomali_student_record.xlsx
  >> SCENARIO (narration):
    Suppose we maintain student academic records and want to automatically flag
    unusual cases that may signal data-entry errors or students needing
    intervention. For each of 200 students we have GPA, weekly study_hour,
    absenteeism_day, test_score, and motivation. A student might look normal on
    any single measure but be highly unusual in the combination of all of them.
    We use Anomaly Detection on this multivariate profile to identify the
    records that deviate most strongly from the typical student pattern.
  >> VARIABLE SELECTION:
    - Features (variables): GPA, study_hour, absenteeism_day, test_score, motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 23   (IQR-based multivariate outlier detection)
    Extraordinary profiles in student records

>> COMMENTARY (narration):
    We found the extraordinary profiles in student records (GPA, study, absenteeism, test, motivation) with anomaly
    detection -- 23 students were flagged as deviating from the normal pattern. These anomalies can point to real
    situations: studying a lot but scoring low (a learning difficulty?), studying little but scoring high (giftedness?),
    or data-entry errors. Anomaly detection automatically picks out the few special cases that need attention from
    large student databases. In education it is extremely practical for early warning, special-education referral and
    data-quality control.

====================================================================================

#100  Variance Components Analysis
    file: 100_varcomp_3seviye_h2.xlsx
  >> SCENARIO (narration):
    Imagine we want to understand where the variability in student achievement
    scores actually comes from in a three-level school system. Our 100
    measurements are nested: each value belongs to a lower_unit (such as a
    class), which is nested within an upper_unit (one of four districts, L_A to
    L_D), with a measurement_id identifying the individual record. Rather than
    testing a single mean difference, we want to know what proportion of the
    total variance lies between districts, between classes, and within classes.
    We use Variance Components Analysis to partition the variance of value
    across these nested levels and estimate, for example, an intraclass measure
    of how much districts and classes matter.
  >> VARIABLE SELECTION:
    - Dependent variable (value): value
    - Random level 1 (upper): upper_unit (L_A/L_B/L_C/L_D)
    - Random level 2 (lower, nested): lower_unit
    - Measurement / observation id: measurement_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper unit = 12.1%   lower unit (nested in upper) = 40.9%   residual = 47.0%
    Nested variance components (REML)

>> COMMENTARY (narration):
    We partitioned the variability of a measurement by which level -- upper unit (e.g. school), lower unit (class) or
    individual -- it comes from with variance-components analysis. The result: 40.9% of the variance is at the lower
    unit (class) level, 12.1% at the upper unit (school) level, and 47% individual/residual variation. So the largest
    structural source is the class level; between-school differences are relatively small. This is a critical
    inference in heritability (h2) and multi-level design studies: it shows where the variability concentrates and at
    which level (class, school or individual) an intervention should be targeted.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old.xlsx
  >> SCENARIO (narration):
    Suppose a school district piloted a new instructional program and wants to
    know whether it raises a student outcome compared with the old program. Each
    of 80 learning units was assigned to either the old or the new program,
    recorded in the group variable, and we measured the resulting outcome in
    value. Instead of a classical t-test that only tells us whether to reject a
    null hypothesis, we run a Bayesian t-test so we can quantify how much more
    probable the difference is under the data and report a Bayes factor and a
    credible interval for the effect. This is the right choice because we want a
    continuous outcome compared across two independent groups while expressing
    our evidence in probabilistic terms.
  >> VARIABLE SELECTION:
    - Grouping factor (two levels): group (old / new)
    - Dependent variable: value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 3.36e+46 (decisive evidence)   Cohen's d = 3.94
    New method vs old method -- score

>> COMMENTARY (narration):
    We tested whether a new teaching method differs from the old with a Bayesian t-test. The Bayes Factor BF10 =
    3.36e+46 -- astronomically large, "decisive evidence": the data support the hypothesis of a difference quadrillions
    of times more than no difference. Cohen's d = 3.94 makes the effect enormous. Unlike a classic p-value, the Bayes
    Factor directly measures the strength of evidence for both H1 and H0 and distinguishes "absence of evidence" from
    "evidence of absence". The new method's effect is not just statistical but practically overwhelming -- the
    Bayesian framework shows this convincingly.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BF10.xlsx
  >> SCENARIO (narration):
    Imagine we are studying whether two continuous educational measures move
    together, for example a study-effort indicator and an achievement indicator
    collected from 80 learners. We have X_variable and Y_variable and want to
    know not just the size of the correlation but how strongly the data support
    the existence of a relationship at all. A Bayesian correlation is
    appropriate here because it returns a posterior distribution for the
    correlation coefficient and a Bayes factor (BF10) that weighs the evidence
    for a non-zero association against the null of independence. This lets us
    say how convincing the link really is rather than relying on a single
    p-value.
  >> VARIABLE SELECTION:
    - First variable (X): X_variable
    - Second variable (Y): Y_variable

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.780 (very strong)   BF10 = 4.50e+14 (decisive evidence)
    Relationship between two variables

>> COMMENTARY (narration):
    We assessed the relationship between two variables with Bayesian correlation: r = 0.78 is very strong and the
    Bayes Factor BF10 = 4.50e+14 makes the evidence for the relationship's existence decisive. Unlike classic
    correlation, Bayesian correlation expresses the strength of the relationship with a probability distribution and
    an evidence factor -- also answering "how sure are we". Such a large BF10 says the probability that the
    relationship is coincidental is vanishingly small. It is a robust framework for quantifying strong links between
    educational variables and reporting the strength of evidence.

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_2yonlu.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether two instructional design factors jointly
    shape a learning outcome. In this study 120 observations were crossed across
    factor1, which has three levels A, B and C, and factor2, which has two
    levels X and Y, and the continuous outcome is value. We use a Bayesian two-
    way ANOVA because we want to compare models with and without each main
    effect and the interaction, and report Bayes factors that tell us which
    combination of factors the data actually favor. This is more informative
    than classical ANOVA when our goal is to weigh competing explanatory models
    rather than simply test significance.
  >> VARIABLE SELECTION:
    - First factor: factor1 (A / B / C)
    - Second factor: factor2 (X / Y)
    - Dependent variable: value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor1: F(2,117) = 17.91, eta2p = 0.234, BF10 = 76,099 (decisive evidence)
    Two-factor design -> score

>> COMMENTARY (narration):
    We examined the effect of two factors on score with Bayesian ANOVA. The first factor's effect is strong: F(2,117)
    = 17.91, effect size eta2p = 0.23, Bayes Factor 76,099 -- "decisive evidence". Bayesian ANOVA gives a separate
    Bayes Factor for each effect and interaction, answering "which factor really matters" more informatively than
    classic ANOVA -- and can even provide "evidence of absence" for non-significant effects. It is a powerful way to
    evaluate factor effects with their strength of evidence in educational experiments (such as method x level).

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_LMM.xlsx
  >> SCENARIO (narration):
    Consider data collected from learners nested within five groups such as
    classrooms or schools, labeled G1 through G5, with 200 observations in
    total. We have a continuous predictor X_covariate and a continuous response
    Y_response, and we suspect that the relationship varies from one group to
    another. A Bayesian hierarchical (multilevel) model is ideal because it lets
    each group have its own intercept and slope while partially pooling
    information across groups, and it gives full posterior distributions for
    both the overall effect and the between-group variability. This respects the
    clustered structure of educational data and avoids treating students as if
    they were independent.
  >> VARIABLE SELECTION:
    - Grouping / random-effect variable: group_id (G1...G5)
    - Predictor (covariate): X_covariate
    - Response (dependent variable): Y_response

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group random variance = 2.90 (8.6%)   residual = 30.79 (91.4%)   ICC = 0.086
    Y ~ X covariate  +  (1 | group)   (Bayesian)

>> COMMENTARY (narration):
    We analysed nested data (observations within groups) with a Bayesian hierarchical model -- the Bayesian
    counterpart of LMM. The group-level variance is only 8.6% of the total (ICC = 0.086); most of the variability
    (91%) is within-group/individual. So in this data between-group differences are small, observations behave
    largely independently. The Bayesian hierarchical model's strength: it expresses both within- and between-group
    uncertainty with full probability distributions and balances groups with few observations via "partial pooling".
    It is a modern method for multi-centre or nested educational data.

====================================================================================

#105  Spatial Lag (SAR) Model
    file: 105_spatial_sar_spatial.xlsx
  >> SCENARIO (narration):
    Suppose we mapped an educational indicator across 100 geographic units such
    as school catchment areas and noticed that neighboring areas tend to have
    similar values. We have the coordinates lat and lon, two area-level
    predictors X1 and X2, and the outcome Y_value. A Spatial Lag (SAR) model is
    appropriate because it adds a spatially lagged term for the outcome,
    capturing the idea that an area's result is influenced by the results of its
    neighbors, which would bias an ordinary regression. This lets us estimate
    the predictor effects while explicitly accounting for spatial spillover
    between schools.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon
    - Predictors: X1, X2
    - Dependent variable: Y_value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho (spatial lag) = 0.34   z = 4.94   p < .001   Pseudo R2 = 0.87
    Y ~ X1 + X2  +  neighbour Y

>> COMMENTARY (narration):
    Modelling a spatial outcome (e.g. regional school achievement), we accounted for spatial dependence -- the
    tendency of nearby units to be similar. The SAR (spatial autoregressive) model links a unit's outcome to its
    neighbours' outcome too (the rho parameter). Here rho = 0.34, significantly positive (z = 4.94, p < .001): the
    neighbour effect is real and moderate -- a region's achievement is influenced by neighbouring regions. The model
    explains 87% of the variance. Using a spatial model is critical, because ordinary regression gives spurious
    significant results when spatial autocorrelation is present. SAR is the right method for spatial educational data
    with diffusion and neighbour effects.

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_residual.xlsx
  >> SCENARIO (narration):
    Imagine the same kind of regional education data over 100 spatial units,
    where the predictors X1 and X2 explain part of an outcome Y_value but the
    leftover residuals are still spatially clustered, perhaps due to unmeasured
    neighborhood factors. The coordinates are stored in lat and lon. Here a
    Spatial Error model is the right tool because the spatial dependence lives
    in the error term rather than in the outcome itself, so we model correlated
    residuals to obtain unbiased standard errors and trustworthy predictor
    estimates. This corrects inference when nearby areas share hidden
    influences.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon
    - Predictors: X1, X2
    - Dependent variable: Y_value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda (spatial error) = 0.59   z = 6.38   p < .001   Pseudo R2 = 0.84
    Y ~ X1 + X2  +  spatial error structure

>> COMMENTARY (narration):
    The spatial error model addresses a different kind of spatial dependence than SAR: not a direct effect of
    neighbour values, but correlation induced through the residuals by unobserved (omitted) spatial factors. Lambda =
    0.59 (z = 6.38, p < .001) shows this spatial error structure is strong -- there are unmeasured common spatial
    factors (regional policy, socio-economic background). Which spatial model is appropriate (SAR or SEM) depends on
    whether the effect comes from neighbour values or from unmeasured factors. SEM cleans out these unmeasured
    spatial confounders to estimate the true effect of the variables more accurately.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local.xlsx
  >> SCENARIO (narration):
    Suppose we suspect that the effect of our predictors on an educational
    outcome is not constant across a region but changes from place to place.
    Across 110 spatial units we have coordinates lat and lon, two predictors X1
    and X2, and the outcome Y_value. Geographically Weighted Regression is
    appropriate because it fits a separate local regression around each
    location, weighting nearby observations more heavily, so we can map how the
    relationship between X1, X2 and Y_value varies geographically. This reveals
    spatial non-stationarity that a single global model would hide.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon
    - Predictors: X1, X2
    - Dependent variable: Y_value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.887   (local/location-specific regression)
    Spatial variation of the relationship

>> COMMENTARY (narration):
    Ordinary regression assumes a single relationship for the whole region; but this relationship can vary in space
    -- a variable's effect may be strong in one region and weak in another. GWR (Geographically Weighted Regression)
    captures this spatial heterogeneity by fitting a separate local regression at each location (R2 = 0.89). The
    output is a map of coefficients: a surface showing where a variable's effect is strong and where it is weak. This
    tests "is the relationship the same everywhere" and sets local policy priorities. In education, GWR reveals what
    the global model hides for mapping regional inequalities and intervention priorities.

====================================================================================

#108  Nested Mixed Model (Nested LMM)
    file: 108_nested_lmm_R_P_F.xlsx
  >> SCENARIO (narration):
    Consider an educational experiment with a strictly hierarchical design: 288
    measurement units are organized so that lower groups labeled F01 through F06
    sit inside upper groups P1 through P4, which in turn sit inside blocks R1
    through R4, and the outcome is value. Because each lower group belongs to
    exactly one upper group and each upper group to one block, the factors are
    nested rather than crossed. A Nested Mixed Model is the correct approach
    because it assigns random effects to each level of nesting, correctly
    partitioning the variance across blocks, upper groups and lower groups while
    estimating the outcome. This avoids treating nested levels as independent
    factors.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Random effect (outermost level): block (R1...R4)
    - Nested random effect (middle level): upper_group (P1...P4)
    - Nested random effect (innermost level): lower_group_no (F01...F06)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper group (P): F = 27.44, p < .001 ***   |   lower x upper (RxP): F = 14.34, p < .001 ***
    block [within upper group] (F): F = 2.20, p = 0.023 *

>> COMMENTARY (narration):
    In this design the measurements are nested: upper group > block > lower unit. The nested mixed model tests each
    level's contribution to the variance separately. The results are significant at all levels: between-upper-group
    differences are largest (F = 27.44), but there are real differences at the block and lower-unit levels too.
    Mixing up the levels in nested data (e.g. ignoring the upper group) produces spurious significance or wrong
    standard errors. Nested LMM is the correct framework for hierarchical sampling designs (school/class/student,
    region/school/class), separating the genuine contribution of each scale.

====================================================================================

#109  Crossed Mixed Model (Crossed LMM)
    file: 109_crossed_lmm_A_B.xlsx
  >> SCENARIO (narration):
    Suppose we ran a study where every level of one factor is observed in
    combination with every level of another, for example each teaching condition
    delivered with each set of materials. Here factor_a has four levels A1 to A4
    and factor_b has four levels B1 to B4, fully crossed across 128
    observations, with the continuous outcome value. A Crossed Mixed Model is
    appropriate because both factors can be treated as random effects that cross
    each other, letting us separate the variability due to factor_a, the
    variability due to factor_b, and the residual, rather than forcing a nested
    structure. This is the standard model for fully crossed educational designs.
  >> VARIABLE SELECTION:
    - Dependent variable: value
    - Crossed random effect 1: factor_a (A1...A4)
    - Crossed random effect 2: factor_b (B1...B4)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor A: F = 0.38, p = 0.77 (non-significant)   |   factor B: F = 0.90, p = 0.48 (non-significant)
    A x B interaction: F = 14.67, p < .001 *** (SIGNIFICANT)

>> COMMENTARY (narration):
    Unlike nesting, here two factors are crossed: each A level is combined with each B level (fully factorial). A
    striking result: both main effects are non-significant (A: p = 0.77, B: p = 0.48), but the interaction is highly
    significant (F = 14.67, p < .001). This is a classic case of a "pure interaction": A and B alone say nothing, but
    TOGETHER -- specific A-B combinations -- they create a strong effect. Looking only at the main effects and saying
    "nothing is significant" would be a major error; the interaction changes everything. The crossed mixed model is
    the right way to capture interactions that emerge when factors are tested jointly -- a critical inference in
    educational experiments (such as method x content).

====================================================================================

#110  KDE Density Map
    file: 110_KDE_school_distribution_density.xlsx
  >> SCENARIO (narration):
    Imagine we have the locations of 95 schools together with an achievement
    indicator, and we want to see where high- and low-performing schools
    concentrate across the map rather than looking at them one point at a time.
    We have coordinates lat and lon and the value school_achievement. A Kernel
    Density Estimation map is the right tool because it smooths the point
    pattern into a continuous density surface, optionally weighted by
    achievement, so we can visually identify clusters and gaps in the
    educational landscape. This turns scattered school points into an intuitive
    heat map for policy discussion.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon
    - Weight (intensity value): school_achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    95 school points   mean achievement = 74.6
    School-distribution density map via kernel density estimation

>> COMMENTARY (narration):
    On the Map tab, we render the geographic density of schools into a heat map with KDE (Kernel Density Estimation).
    KDE turns points into a smooth density surface: we can see where school clustering is dense and where there are
    gaps. The distribution of 95 schools shows how educational infrastructure concentrates geographically. This is
    used both to detect access gaps (which regions have few schools) and to see dense regions. It is the basic
    mapping tool for visualising spatial point data; in educational planning it assesses the equity of school
    distribution.

====================================================================================

#111  Hexbin Map
    file: 111_Hexbin_student_distribution.xlsx
  >> SCENARIO (narration):
    Suppose we have geocoded the home locations of 150 students across a
    metropolitan area and we want to see where students concentrate spatially
    rather than just plotting hundreds of overlapping points. Each student
    record carries a latitude and longitude, plus a density_interval label
    classifying the surrounding area as low, medium, or high. A Hexbin Map is
    ideal here because it bins thousands of point coordinates into uniform
    hexagonal cells and shades each cell by the count, turning a cluttered dot
    map into a clean density surface. This lets us instantly spot the
    neighbourhoods where the student population clusters and where coverage is
    sparse, which is exactly what we need for planning school catchment areas.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon
    - Weight / category (optional): density_interval (low / medium / high)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    150 student/unit points   aggregated into hexagonal cells
    Each cell shows the point density within it by colour

>> COMMENTARY (narration):
    The hexbin map aggregates many points into hexagon cells and shows each cell's density by colour -- a clear
    solution, alternative to KDE, when thousands of points overlap into an unreadable mess. We rendered 150 student/
    unit points into a hexagonal grid to visualise where density is highest. Hexagonal cells carry less edge-bias
    than squares and represent neighbour relations more evenly. It is the practical way to turn dense point data
    (student distribution, school location) into a readable density map.

====================================================================================

#112  Moran's I (Spatial Autocorrelation)
    file: 112_Morans_I_achievement_autocorrelation.xlsx
  >> SCENARIO (narration):
    Imagine we have school-level achievement scores for 80 schools spread across
    a region, each with its latitude and longitude. The research question is
    whether high-achieving schools tend to sit next to other high-achieving
    schools, or whether achievement is scattered randomly across space. Moran's
    I is the right tool because it is the classic global measure of spatial
    autocorrelation: it tests whether the value of school_achievement at one
    location is correlated with the values at neighbouring locations. A
    significant positive Moran's I would tell us that educational achievement is
    spatially clustered, which has direct implications for resource allocation
    and equity policy.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon
    - Value variable: school_achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.530   z = 9.40   p < .001
    Spatial autocorrelation of school achievement (strong positive)

>> COMMENTARY (narration):
    Moran's I is the global spatial-autocorrelation statistic measuring "do nearby schools have similar achievement".
    The result is strongly positive (I = 0.53, z = 9.40, p < .001): school achievement is spatially clustered --
    high-achieving schools lie near each other, and so do low ones. This is not chance; shared region, socio-economic
    environment and resource distribution make neighbouring schools similar. This finding is a concrete indicator of
    spatial inequality in education and is methodologically important: if spatial autocorrelation exists, ordinary
    statistics are wrong (hence the spatial models in #105-107). Moran's I is the starting diagnostic of spatial analysis.

====================================================================================

#113  Getis-Ord Hotspot
    file: 113_Getis_Ord_high_achievement_hotspot.xlsx
  >> SCENARIO (narration):
    Now suppose we don't just want a single global statistic, but we want to
    pinpoint exactly where the clusters of high and low achievement are located
    among our 68 schools. Each school has coordinates and a school_achievement
    score. The Getis-Ord Gi* hotspot statistic is designed precisely for this:
    it scans each location together with its neighbours and flags statistically
    significant hot spots, where high achievement piles up, and cold spots,
    where low achievement concentrates. This gives education planners a map of
    priority intervention zones rather than just an overall summary number.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon
    - Value variable: school_achievement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    29 hot spots (high-achievement cluster)   31 cold spots (low-achievement cluster)   n = 68
    Getis-Ord Gi* local statistic (z > 1.96 significant)

>> COMMENTARY (narration):
    While Moran's I reports the overall clustering of the whole region, Getis-Ord Gi* shows WHERE the hot/cold spots
    are, location by location. 29 schools are statistically significant "hot spots" (high achievement surrounded by
    high achievement), and 31 schools are "cold spots" (low-achievement clusters). This gives direct guidance for
    education policy: it is most efficient to focus resources and intervention on the cold-spot clusters
    (disadvantaged regions). Gi* hot-spot analysis is the standard spatial method for finding "where the problem/
    success concentrates geographically" -- it maps educational inequality.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_school_kumeleri.xlsx
  >> SCENARIO (narration):
    Finally, suppose we have the geographic coordinates of 115 schools and we
    want to discover natural geographic clusters of schools without deciding in
    advance how many clusters there are. Each record holds only a latitude and a
    longitude. DBSCAN is well suited to this because it is a density-based
    clustering method: it groups schools that lie close together into clusters
    and automatically labels isolated, far-flung schools as noise, all without
    us pre-specifying the number of clusters. This helps us identify dense
    school networks versus remote, under-served locations.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   noise (isolated) = 6   n = 115
    School locations clustered by density

>> COMMENTARY (narration):
    On the map we clustered schools' geographic locations by density with DBSCAN: three dense school groups were
    found, and 6 schools were flagged as "isolated/noise" points belonging to no cluster. DBSCAN's strength is that
    it does not require the number of clusters in advance and can capture irregularly shaped clusters -- and
    automatically separates sparse/distant schools. The three clusters show schools group in particular city centres/
    regions; the isolated schools may be rural/remote units. This spatial clustering is a practical tool for defining
    educational service regions, planning resource distribution and assessing geographically similar schools
    together. With this we complete the full 114-analysis tour of the Educational Sciences package.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_achievement_score.xlsx
  >> SCENARIO (narration):
    We follow 40 students measured at three test levels (pretest, posttest, retention); each belongs to one of two method groups (experimental / control). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on achievement score.
  >> VARIABLE SELECTION:
    - Dependent variable: achievement_score
    - Subject ID: student_id
    - Between-subjects factor: method
    - Within-subjects factor: test

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (method): F(1,38) = 7.91  p = 0.008  np2 = 0.172
    Within (test):  F(2,76) = 157.82  p < .001  np2 = 0.806
    Interaction:         F(2,76) = 17.95  p < .001  np2 = 0.321
    Mauchly W = 0.928  p = 0.244   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a method x test mixed design we analyzed achievement score for 40 students (120 observations). The interaction is significant (F(2,76) = 17.95, p < .001, np2 = 0.321) *** -- the two groups' change across test differs in magnitude. The between-subjects main effect (experimental vs control) is F = 7.91, p = 0.008; the within-subjects main effect (pretest/posttest/retention) is F = 157.82, p < .001. Mauchly's test p = 0.244, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each test level. In educational sciences, the mixed design is the standard analysis for pretest-posttest-retention designs comparing teaching methods.

====================================================================================
