====================================================================================
MerQur - Psikoloji_Sosyoloji - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Psikoloji_Sosyoloji/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_demographics.xlsx
  >> SCENARIO (narration):
    Imagine we have just collected survey data from 280 community members and we
    want a clear first picture of who they are and how they live. The dataset
    records each person's occupational type (participant, worker, retired,
    employer, or civil servant), their age, years of education, monthly income
    in Turkish Lira, a life-satisfaction rating, and how many hours of weekly
    activity they report. Before running any inferential test, we want to
    summarize these variables: typical values, spread, and how the groups are
    distributed. Descriptive statistics is the right starting point here because
    we are not testing a hypothesis yet, we simply want means, standard
    deviations, ranges, and frequency counts to understand the sample and spot
    any anomalies.
  >> VARIABLE SELECTION:
    - Numeric variables: age
    - Numeric variables: education_year
    - Numeric variables: monthly_income_TL
    - Numeric variables: satisfaction
    - Numeric variables: weekly_activity
    - Grouping (categorical) variable: type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 participants
    age      : mean = 43.9        |   education years : mean = 13.1        |   monthly income : mean = 12,228 TL
    life satisfaction : mean = 3.47

>> COMMENTARY (narration):
    We first drew the general socio-demographic profile of our participant sample: 280 participants, mean age 44, 13
    years of education, mean monthly income 12,228 TL, life satisfaction 3.47 (moderate-good on a 5-point scale). This
    descriptive table sets the stage for all the analyses that follow -- group comparisons, well-being relationships,
    income-satisfaction links. In the social sciences, before any inferential test, summarising the sample's basic
    demographic characteristics is essential both to check data quality and to place the findings in the right
    context.

====================================================================================

#2  Normality Tests
    file: 02_normality_income_motivation.xlsx
  >> SCENARIO (narration):
    Suppose we are studying how income relates to psychological states in a
    sample of 200 adults, and we have measured each person's income, motivation,
    and anxiety. Many of the parametric tests we plan to run later, like t-tests
    and ANOVA, assume that these continuous variables are approximately normally
    distributed. Before trusting those tests, we want to check that assumption
    directly. Normality tests, such as Shapiro-Wilk together with skewness and
    kurtosis, are appropriate because they formally evaluate whether each
    variable departs from a normal distribution, which guides whether we use
    parametric or nonparametric methods.
  >> VARIABLE SELECTION:
    - Test variables: income
    - Test variables: motivation
    - Test variables: anxiety

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    income     : Shapiro-Wilk = 0.899  p < .001    KS p = 0.001   Not Normal
    motivation : Shapiro-Wilk = 0.994  p = 0.657   KS p = 0.636   Normal
    anxiety    : Shapiro-Wilk = 0.904  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three variables -- income, motivation and anxiety -- are normally distributed. The result
    differs by variable: motivation is normal (p = 0.657 > 0.05), but income and anxiety deviate significantly from
    normality (p < .001). This is very typical in the social sciences: income is almost always right-skewed (most
    people low-to-middle income, a few very high), and anxiety scales can be skewed too. The practical takeaway: we
    can safely use parametric tests (t-test, ANOVA) on motivation; for income and anxiety, nonparametric methods
    (Mann-Whitney, Kruskal-Wallis) or a log-transform are more appropriate. The normality check is a critical pre-step
    that decides, for each variable separately, which test family is appropriate.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_wellbeing.xlsx
  >> SCENARIO (narration):
    Imagine a community well-being program that defines a target benchmark
    score, and we want to know whether the public community we sampled meets
    that standard. We surveyed 100 members of the public and measured each
    person's well-being on a continuous scale. Our question is simple: does the
    average well-being of this group differ from a known reference value, say
    the policy benchmark? A one-sample t-test is the correct choice because we
    are comparing the mean of a single continuous variable against a fixed,
    theoretical value rather than against another group.
  >> VARIABLE SELECTION:
    - Test variable: wellbeing
    - Test value (mu): policy benchmark constant (e.g. 50)

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = 2.3164   p = 0.023 *   Cohen d = 0.232 (Small)
    Mean well-being = 54.80   (test mu = 50, scale midpoint)   H0 REJECTED

>> COMMENTARY (narration):
    We compared a community's mean well-being score against the scale midpoint of 50. The result: mean 54.80,
    significantly above the midpoint -- t(99) = 2.32, p = 0.023, but the effect size is small at d = 0.23. So the
    community is slightly above the midpoint, but the difference is modest. This distinction matters: a small effect
    reminds us that statistical significance is not the same as practical importance. The one-sample t-test is the
    right way to compare a group mean against a known reference (scale midpoint, norm value) -- common in
    psychological scale research.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent Samples t-Test
    file: 04_independent_t_sex_score.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether men and women differ in their performance on
    a psychological assessment. In this dataset 115 participants were classified
    by sex as female or male, and each completed the same test, giving a
    continuous score. The research question is whether the average score differs
    between the two sex groups. An independent samples t-test fits perfectly
    here because we are comparing the means of one continuous outcome across two
    separate, unrelated groups.
  >> VARIABLE SELECTION:
    - Grouping variable: sex (female, male)
    - Dependent variable: score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = 3.606   p < .001 ***   Cohen d = 0.673 (Moderate)
    Score differs significantly by sex.   H0 REJECTED

>> COMMENTARY (narration):
    We examined whether a psychological scale score differs by sex with an independent t-test. The difference is
    significant and moderate (t(113) = 3.61, p < .001, d = 0.67): the two sexes' scores differ markedly. The moderate
    effect size shows the difference is both statistically and practically notable. The independent t-test is the
    basic method for comparing the means of two independent groups (sex, community, group); reporting effect size
    alongside the p-value is standard in the social sciences.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired t-Test
    file: 05_paired_t_on_last.xlsx
  >> SCENARIO (narration):
    Imagine an intervention designed to change participants' performance, where
    each of the 50 people was tested before and after taking part. We recorded
    an initial measurement, the on_test score, and a later measurement, the
    last_test score, for the very same individuals. The question is whether
    scores changed from the first assessment to the second. Because the two
    measurements come from the same people, they are paired, so a paired-samples
    t-test is appropriate to test whether the mean difference between the two
    time points is significantly different from zero.
  >> VARIABLE SELECTION:
    - Pair measurement 1: on_test
    - Pair measurement 2: last_test

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Paired t(49) = -15.99   p < .001 ***   Cohen d_z = -2.262 (Large)
    Mean difference (pre-test − post-test) = -12.23   H0 REJECTED

>> COMMENTARY (narration):
    We measured an intervention's (therapy, program) effect by comparing the same participants' pre- and post-test
    scores with a paired t-test. The result is striking: post-test scores are on average 12.23 points higher than
    pre-test, t(49) = -15.99, p < .001, with an enormous effect size (d_z = -2.26). The intervention produced a clear
    improvement. Because we measured the same participants twice, the paired test is correct -- it removes
    between-person variability and focuses only on the individual's own change. It is the gold-standard way to
    evaluate before/after interventions in psychology.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_anova_therapy.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether different counseling and instructional
    approaches change people's outcomes. In this dataset 140 participants each
    received one of four therapy types: standard, cooperative, project, or
    mixed, and afterward we measured a continuous outcome score. The research
    question is whether the average score differs across these four therapy
    groups. We use one-way ANOVA because the outcome is continuous and the
    predictor is a single categorical factor with more than two levels, allowing
    us to test for any overall group differences before looking at specific
    pairs.
  >> VARIABLE SELECTION:
    - Factor (group): therapy (standard, cooperative, project, mixed)
    - Dependent variable: score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 136) = 15.3770   p < .001 ***   η² = 0.2533
    DECISION: H0 REJECTED (4 therapy types)

>> COMMENTARY (narration):
    We compared the effect of four therapy types on score with one-way ANOVA. The result is significant (F(3,136) =
    15.38, p < .001), with effect size η² = 0.25 -- a quarter of the score is explained by therapy type, a strong
    effect in clinical psychology. ANOVA says at least one therapy differs from the others; a post-hoc test is needed
    to see which. "Which therapy/intervention is more effective" is a fundamental research question in clinical
    psychology; ANOVA is the standard way to answer it across multiple groups.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova.xlsx
  >> SCENARIO (narration):
    Imagine we are studying how a teaching method and a person's sex jointly
    affect performance. Here 150 participants were assigned to one of three
    methods, standard, project, or mixed, and they are also split by sex into
    female and male, with a continuous score recorded for each. We want to know
    not only whether method matters and whether sex matters, but also whether
    the effect of method depends on sex. A two-way ANOVA is ideal because it
    lets us test both main effects and, crucially, the interaction between the
    two categorical factors on the continuous outcome.
  >> VARIABLE SELECTION:
    - Factor 1: method (standard, project, mixed)
    - Factor 2: sex (female, male)
    - Dependent variable: score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    method (main) : F(2) = 43.85   p < .001 ***   η²p = 0.38
    sex (main)    : F(1) =  0.42   p = 0.519 ns    η²p = 0.003

>> COMMENTARY (narration):
    We examined score with two factors -- method and sex -- together. The method effect is very strong (F = 43.85,
    η²p = 0.38), while the sex effect is non-significant (F = 0.42, p = 0.519). So the only determinant of the outcome
    is method; sex has no effect. A non-significant sex effect is information too: the intervention works similarly
    for both sexes, so no sex-specific adaptation is needed. Two-way ANOVA's strength is seeing both factors' effects
    -- separately and jointly (interaction) -- at once; it answers "does the method effect vary by sex".

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated Measures ANOVA
    file: 08_repeated_anova_wave_wellbeing.xlsx
  >> SCENARIO (narration):
    Suppose we are tracking the well-being of 60 individuals across a program
    and we measured each person at four successive time points. The dataset
    stores these as measurement_1 through measurement_4, all from the same
    units. We want to know whether well-being changes systematically across the
    four waves of measurement. A repeated measures ANOVA is the right tool
    because the four scores are repeated observations on the same subjects, so
    it accounts for the within-person correlation while testing for differences
    across time points.
  >> VARIABLE SELECTION:
    - Repeated measures (within-subject): measurement_1
    - Repeated measures (within-subject): measurement_2
    - Repeated measures (within-subject): measurement_3
    - Repeated measures (within-subject): measurement_4

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 177) = 108.40   p < .001 ***   η²p = 0.6475   (n = 60)
    DECISION: H0 REJECTED

>> COMMENTARY (narration):
    We compared the same 60 participants' well-being scores across four consecutive waves with repeated-measures
    ANOVA. Because the same people are measured repeatedly, observations are dependent; RM-ANOVA accounts for this
    within-subject correlation. The result is very strong (F(3,177) = 108.40, p < .001, η²p = 0.65): well-being
    changes significantly across waves -- a strong temporal trend. RM-ANOVA is the correct design for longitudinal
    (panel) studies where the same people are tracked over time, and is far more powerful than treating each wave as
    independent. It is ideal for panel research following social change.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_method_3domain.xlsx
  >> SCENARIO (narration):
    Imagine we want to know whether a teaching method affects a whole profile of
    abilities rather than a single outcome. In this study 120 participants were
    assigned to one of three methods, standard, mixed, or project, and we
    measured three correlated outcomes for each: math, science, and language
    performance. Because these three outcomes likely move together, testing them
    separately would inflate error and miss their joint pattern. MANOVA is
    appropriate here because it tests whether the group means differ across the
    combined set of multiple continuous dependent variables simultaneously.
  >> VARIABLE SELECTION:
    - Factor (group): method (standard, mixed, project)
    - Dependent variables: math
    - Dependent variables: science
    - Dependent variables: language

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.433 -> F(6, 230) = 19.90   p < .001 ***   (n = 120)
    Dependent: 3 domains (math, science, language)   |   Factor: method

>> COMMENTARY (narration):
    Here we have three correlated dependent variables (three skill domains) and one method factor. Running three
    separate ANOVAs would inflate the error rate and ignore the correlations among domains. MANOVA tests all three
    jointly: Wilks' Lambda F = 19.90, p < .001, so the method strongly shifts the multivariate profile. MANOVA is the
    right approach when an intervention affects several related outcomes at once -- it controls the overall error
    rate. After a significant MANOVA, follow-up univariate tests show which domain drives the effect.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova.xlsx
  >> SCENARIO (narration):
    Suppose we are evaluating an intervention with a control group and two
    treatment arms, intervention-A and intervention-B, across 105 participants.
    Everyone completed a pre-test, the on_test score, and a post-test, the
    last_test score. We want to compare the post-test outcomes across the three
    groups, but it would be unfair to ignore that people started at different
    levels. ANCOVA is the right method because it compares the adjusted group
    means on the post-test while statistically controlling for the pre-test as a
    covariate, isolating the effect of group above and beyond baseline
    differences.
  >> VARIABLE SELECTION:
    - Factor (group): group (control, intervention-A, intervention-B)
    - Dependent variable: last_test
    - Covariate: on_test

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    group (factor): η²p = 0.474 (large)   p < .001 ***
    Covariate: pre-test (baseline level)   |   DV: post-test   H0 REJECTED

>> COMMENTARY (narration):
    When evaluating an intervention's effect, it is essential to control for participants' baseline (pre-test) levels
    -- because if the groups differ at baseline, the post-test difference is misleading. ANCOVA takes the pre-test as
    a covariate and holds it statistically constant. The result: even after adjusting for baseline, the groups differ
    significantly in post-test (η²p = 0.47, large effect). So the observed difference is not a by-product of baseline
    differences; it is the genuine intervention effect. ANCOVA is the standard way -- very common in psychology -- to
    fairly control baseline differences in pre/post intervention studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap Confidence Interval
    file: 11_bootstrap_ci.xlsx
  >> SCENARIO (narration):
    Imagine we are studying how long people stay engaged before withdrawing from
    a community wellbeing program. In this dataset 30 participants from public
    and private communities were tracked, and we recorded the withdrawal_day for
    each person. With such a small sample we are reluctant to assume the
    withdrawal times follow a neat normal distribution, so instead of a textbook
    formula we resample the data thousands of times. A bootstrap confidence
    interval lets us estimate a stable range for the mean withdrawal day
    directly from the data, which is exactly what we want when the sample is
    small and the shape of the distribution is uncertain.
  >> VARIABLE SELECTION:
    - Variable: withdrawal_day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed mean (withdrawal day) = 4.53
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    We derived a confidence interval for the mean of withdrawal (social withdrawal) days by bootstrapping: the
    observed mean is 4.53. Such count/behavioral data are typically right-skewed, and classic normal-assuming formulas
    can be unreliable for such data. Bootstrap estimates the interval without distributional assumptions, through
    thousands of resamples. It is a robust way to express the uncertainty of a mean in skewed psychosocial data
    (withdrawal, behavior counts).

====================================================================================

#12  Permutation Test
    file: 12_permutation.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether a motivational intervention actually raises
    motivation. Here 53 participants were split into a control group and an
    experiment group, and we measured each person's motivation afterward. Rather
    than relying on the assumptions of a classic t-test, we shuffle the group
    labels many times to build the distribution of mean differences we would
    expect purely by chance. A permutation test is ideal here because it makes
    almost no distributional assumptions and tells us how unusual the observed
    difference between control and experiment really is.
  >> VARIABLE SELECTION:
    - Group (factor): group
    - Test variable: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = -2.25   p = 0.029 *   H0 REJECTED
    Motivation differs significantly between the two groups

>> COMMENTARY (narration):
    We tested the motivation difference between two groups with a permutation test. Permutation shuffles the group
    labels thousands of times to measure whether the observed difference could arise by chance -- requiring no
    distributional assumption. The result is significant (p = 0.029): the groups' motivation differs more than chance
    allows. For small samples or when normality is doubtful, permutation is a robust, assumption-free alternative to
    the t-test. It is a reliable backup for group comparisons in psychological experiments.

====================================================================================

#13  Multiple Comparison Corrections
    file: 13_multiple_comparison.xlsx
  >> SCENARIO (narration):
    Picture an experiment comparing nine different learning materials, eight new
    versions labelled M1 through M8 plus a control, on participants' scores.
    With 180 participants spread across these nine conditions, running every
    pairwise comparison inflates the chance of a false positive dramatically.
    That is why we apply multiple comparison corrections: after testing the
    score differences across the material groups, methods like Bonferroni or
    Holm adjust the p-values so our family-wise error rate stays controlled
    while we still see which materials genuinely differ from the control.
  >> VARIABLE SELECTION:
    - Group (factor): material
    - Dependent variable: score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Raw (uncorrected): 5/8 significant   |   After correction (Bonferroni/Holm/FDR): 3/8 significant

>> COMMENTARY (narration):
    When we compare different materials pairwise on score, we run many tests -- each carrying a false-positive risk.
    Multiple-comparison correction controls that risk. Here is an important lesson: before correction 5 of 8
    comparisons looked significant, but after correction only 3 remain. So 2 "significant" results were actually
    false positives. This shows very clearly why correction is critical when running multiple tests; skipping it
    leads to reporting spurious differences -- a contributor to the replication crisis in the psychology literature.

====================================================================================

#14  Mann-Whitney U Test
    file: 14_mann_whitney.xlsx
  >> SCENARIO (narration):
    Let's say we are comparing life satisfaction between people living in urban
    versus rural locations. In this study 85 participants reported their
    survival_satisfaction, and the two groups are urban and rural. Satisfaction
    scores like these are often skewed and not normally distributed, so
    comparing their means with a t-test would be questionable. The Mann-Whitney
    U test compares the rank distributions of the two independent groups
    instead, making it the right choice for an ordinal-style outcome with an
    uncertain shape.
  >> VARIABLE SELECTION:
    - Group (factor): location
    - Test variable: survival_satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Mann-Whitney U = 1121.50   p = 0.049 *   r = -0.25
    Life satisfaction differs significantly by location (borderline).   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether life satisfaction differs by location (urban/rural) with the nonparametric Mann-Whitney --
    because satisfaction scores are ordinal and may not be normal. The result is borderline significant (U = 1121.5,
    p = 0.049), with a small-moderate effect size r = 0.25. The p-value is very close to 0.05 -- so the evidence is
    weak and should be interpreted cautiously. For ordinal/skewed satisfaction data Mann-Whitney is more reliable than
    the t-test and is widely used in quality-of-life/attitude studies. Borderline results should be confirmed with a
    larger sample.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank Test
    file: 15_wilcoxon_attitude.xlsx
  >> SCENARIO (narration):
    Consider a training study where 35 counselors completed an attitude
    questionnaire before and after an education program. Each counselor gives us
    a paired measurement, attitude_before and attitude_post, so we are
    interested in the within-person change. Because attitude scores are not
    guaranteed to be normally distributed and the sample is modest, we avoid the
    paired t-test. The Wilcoxon signed-rank test ranks the paired differences
    and tests whether attitudes shifted systematically after the program, which
    fits this matched before-and-after design perfectly.
  >> VARIABLE SELECTION:
    - Pair M1 (before): attitude_before
    - Pair M2 (after): attitude_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilcoxon W = 0.00   p < .001 ***   r = 0.891   n (nonzero) = 30
    Counselor attitude changed significantly before vs after training.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same counselors' attitude scores before and after a training with the paired Wilcoxon test. The
    result is very strong: W = 0, p < .001, effect size r = 0.89. Attitudes changed in the same direction and by a
    large amount in nearly all counselors -- the training had a clear effect. For paired measures and ordinal/skewed
    attitude data, Wilcoxon is the right choice. It is a robust, assumption-free way to measure the effect of
    professional-development interventions.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis Test
    file: 16_kruskal_sound_book.xlsx
  >> SCENARIO (narration):
    Imagine we want to know whether background sound level affects how many
    books people read. In this dataset 45 participants were placed under low,
    medium, or high sound conditions, and we counted each person's book_count.
    Book counts are discrete and clearly skewed, and the groups are small, so a
    one-way ANOVA is not appropriate. The Kruskal-Wallis test compares the rank
    distributions across the three independent sound groups, which is the proper
    nonparametric way to test for differences among more than two groups.
  >> VARIABLE SELECTION:
    - Group (factor): sound
    - Test variable: book_count

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 15.96   p < .001 ***   eta-squared_H = 0.33
    Book-reading count differs significantly by socio-economic level.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the number of books read varies by socio-economic level with Kruskal-Wallis -- the
    nonparametric ANOVA for more than two groups, appropriate because count data are not normally distributed. The
    result is highly significant (H(2) = 15.96, p < .001), with a large effect (eta-squared = 0.33): reading is
    strongly tied to socio-economic level. This is a concrete indicator of debates on cultural capital and
    inequality. For count/ordinal data Kruskal-Wallis is more appropriate than classic ANOVA.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman Test
    file: 17_friedman_program.xlsx
  >> SCENARIO (narration):
    Suppose every participant tried four different program conditions and we
    want to know whether they perform differently across them. Here 30 units
    were each measured under condition_A, condition_B, condition_C, and
    condition_D, so all four scores come from the same person, a repeated-
    measures layout. Since we cannot assume normality and the data are matched
    within each unit, a repeated-measures ANOVA is risky. The Friedman test
    ranks the four conditions within each participant and tests whether the
    conditions differ overall, which is exactly the nonparametric analogue we
    need.
  >> VARIABLE SELECTION:
    - Repeated measures (items): condition_A
    - Repeated measures (items): condition_B
    - Repeated measures (items): condition_C
    - Repeated measures (items): condition_D

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Friedman chi-square(3) = 39.08   p < .001 ***   Kendall W = 0.4342   n = 30
    The same participants' scores under 4 conditions differ significantly.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same participants' scores under four different conditions/programs with the Friedman test -- the
    nonparametric counterpart of repeated-measures ANOVA. The result is highly significant (chi-square(3) = 39.08,
    p < .001), with Kendall W = 0.43 indicating moderate-high consistency of the ranking across conditions. So the
    four conditions make a clear difference in participant performance. Because the same people are measured
    repeatedly, observations are dependent; Friedman accounts for this correctly. It is the right test for ordinal/
    non-normal repeated-measures data.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_attainment.xlsx
  >> SCENARIO (narration):
    Let's say a program claims that a certain proportion of participants reach
    an attainment target. In this large cohort of 800 participants we have a
    simple pass indicator, the passed variable, coded as success or failure for
    each person. We want to test whether the observed success rate differs from
    a specific expected proportion, for example fifty percent. The binomial test
    is the natural tool here because the outcome is a single yes-or-no event
    repeated across independent participants, and it gives an exact probability
    rather than relying on a large-sample approximation.
  >> VARIABLE SELECTION:
    - Outcome (binary): passed

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed attainment proportion = 0.7037   test proportion = 0.50
    p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether an attainment/success proportion equals chance (50%) with the binomial test. The observed
    proportion is 0.70 -- about three-quarters of participants reached attainment. This differs significantly from
    50% (p < .001). The binomial test is the simplest way to compare a single binary proportion (success/failure,
    present/absent) against a known reference. In the social sciences it is a practical tool for comparing success/
    attainment rates against a target or chance level.

====================================================================================

#19  Sign Test
    file: 19_sign_test_program.xlsx
  >> SCENARIO (narration):
    Picture a small pilot program where 28 participants were assessed before and
    after, giving us score_before and score_post for each person. We simply want
    to know whether the program tended to push scores up or down, without
    trusting the exact size of each change. The sign test ignores the magnitude
    of the differences and just counts how many participants improved versus
    declined, testing whether positive and negative changes are balanced. With
    such a small paired sample and no assumption about the shape of the
    differences, this is the most robust choice.
  >> VARIABLE SELECTION:
    - Pair M1 (before): score_before
    - Pair M2 (after): score_post

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive (before > after) = 0   (scores INCREASED in ALL participants)   p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same participants' scores before and after a program with the paired sign test. A striking
    result: no participant's score decreased -- all increased (0 positive, i.e. the before-score was higher in none),
    p < .001. So the program raised scores in every single participant. The sign test looks only at the direction of
    change (up/down), not its magnitude; it is a robust, assumption-free test. When the direction of change (how many
    improved) matters, or when the distribution is very skewed, it is a coarse but sturdy before/after indicator.

====================================================================================

#20  Runs Test
    file: 20_runs_test.xlsx
  >> SCENARIO (narration):
    Imagine we are monitoring daily absence over 60 consecutive days to see
    whether absences occur randomly or in clustered streaks. The data record,
    for each day, whether someone was absent, a binary sequence ordered in time.
    We are not asking how many absences there were but whether they cluster into
    runs, which could signal an underlying pattern like contagion or weekly
    cycles. The runs test examines the sequence of outcomes and tests the null
    hypothesis of randomness, making it the right tool for detecting non-random
    ordering in this time-ordered binary series.
  >> VARIABLE SELECTION:
    - Sequence (binary, ordered): absent
    - Order: day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 28   Z = -0.24   (accepted as random)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the daily absence series is randomly distributed -- whether successive days cluster or show a
    pattern -- with the Wald-Wolfowitz runs test. With Z = -0.24 the observed number of runs is very close to
    expected; the series is accepted as random, with no systematic pattern. This matters methodologically: if a
    series is non-random (autocorrelated), many classic tests become invalid. Because randomness holds here, standard
    methods can be used safely. The runs test is ideal for this pre-check on time-ordered behavioral data.

====================================================================================

#21  Chi-Square Test of Independence
    file: 21_chisquare_sound_institution.xlsx
  >> SCENARIO (narration):
    Suppose we are interested in whether the level of perceived environmental
    noise people are exposed to is related to the type of institution they
    belong to. In this dataset of 540 participants, we recorded the perceived
    sound level as low, medium, or high, and classified each person's
    institution as public or private. Our research question is whether noise
    exposure is distributed independently of institution type, or whether one
    sector tends to report louder environments than the other. Because both
    variables are categorical and we want to test whether their joint
    distribution departs from independence, the Chi-Square Test of Independence
    is the appropriate choice. A significant result would tell us that sound
    level and institution type are statistically associated.
  >> VARIABLE SELECTION:
    - Row variable: sound
    - Column variable: institution

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(2) = 141.06   p < .001 ***   Cramer V = 0.51 (Strong)
    Socio-economic level is associated with institution preference.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether socio-economic level is associated with the preferred institution using chi-square. The result
    is very strong: chi-square(2) = 141.06, p < .001, Cramer V = 0.51. So the two variables are not independent --
    socio-economic level markedly shapes institution preference. This is an important indicator of opportunity
    inequality and social stratification in sociology. For association between two categorical variables the
    chi-square test of independence is the standard method; Cramer V measures the strength.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_style.xlsx
  >> SCENARIO (narration):
    Imagine a learning-styles survey where we asked 200 participants to identify
    their dominant style, classified as kinesthetic, auditory, visual, or
    enrolled. We want to know whether these four styles occur equally often in
    the population, or whether some styles are clearly more common than others.
    There is no second variable here, only the observed frequencies of a single
    categorical variable compared against an expected even distribution. Because
    we are testing whether one set of observed category counts matches a
    hypothesized distribution, the Chi-Square Goodness-of-Fit Test is exactly
    what we need. A significant chi-square would indicate that the learning
    styles are not uniformly distributed.
  >> VARIABLE SELECTION:
    - Category variable: style

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 22.88   p < .001 ***   n = 200   k = 4 styles (expected: equal 1/k)
    H0 REJECTED

>> COMMENTARY (narration):
    We tested whether participants' styles (e.g. attachment/cognitive styles) are equally distributed with the
    chi-square goodness-of-fit test. The observed distribution deviates significantly from the equal expectation
    (chi-square(3) = 22.88, p < .001): some styles are more common than expected, others rare. So styles are not
    balanced in the population. The goodness-of-fit test is the right way to compare an observed category distribution
    against a theoretical expectation (equal or specific proportions); it is used in psychological typology and
    population-distribution studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher.xlsx
  >> SCENARIO (narration):
    Consider a small pilot study comparing two counseling methods, standard and
    mixed, to see whether they differ in producing a successful outcome, coded
    as pass yes or no. We only have 28 participants, so the cell counts in our
    two-by-two table are small. With such limited data, the usual chi-square
    approximation can be unreliable. That is why we turn to Fisher's Exact Test,
    which computes the exact probability of the observed association without
    relying on large-sample assumptions. It lets us ask, even with few cases,
    whether the pass rate genuinely depends on the method used.
  >> VARIABLE SELECTION:
    - Row variable: method
    - Column variable: pass

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher Exact p = 0.0178 *   Phi = 0.49 (Strong)
    Method is associated with passing/success.   H0 REJECTED

>> COMMENTARY (narration):
    We tested the association between method and passing/success with Fisher's exact test -- because chi-square is
    unreliable for tables with few observations (small cells). The result is significant (p = 0.018), with Phi = 0.49
    indicating a strong association (the odds ratio took an extreme value because one cell is empty -- a strong,
    near-complete separation). Fisher's exact test gives more accurate results than chi-square in small-sample pilot
    studies or with rare outcomes (small-cell 2x2 tables). It is often needed in small-sample social-science
    experiments.

====================================================================================

#24  McNemar Test
    file: 24_mcnemar.xlsx
  >> SCENARIO (narration):
    Suppose we ran a before-and-after intervention with 90 participants and
    recorded each person's binary status twice: once on an initial test and once
    on a later test, with categories Bsr and Bsm. We are not interested in
    whether the two categories are equally frequent overall, but in whether the
    intervention shifted people from one category to the other, that is, whether
    the pattern of change is symmetric. Because we have paired categorical
    measurements on the same individuals at two time points, the McNemar Test is
    appropriate. It focuses on the discordant pairs to tell us whether there was
    a real change in classification from the first test to the second.
  >> VARIABLE SELECTION:
    - First measurement (pair 1): on_test
    - Second measurement (pair 2): last_test

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar chi-square = 0.00   p = 1.000   discordant: b = 13, c = 14
    H0 NOT REJECTED (no significant change)

>> COMMENTARY (narration):
    We compared the same participants' pre- and post-test pass status (pass/fail binary) with McNemar's test. The
    result is non-significant (chi-square = 0, p = 1.0): 13 people changed positively, 14 negatively -- the balance
    is almost exactly even, so there is no significant shift in overall pass status. A non-significant result is
    informative too: this intervention did not change the binary pass status (it may have affected scores without
    pushing them over the threshold). McNemar looks only at the cells that changed (discordant cells); it is the
    right way to measure before/after categorical change.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa.xlsx
  >> SCENARIO (narration):
    Imagine two counselors independently rated the same 70 clients on an ordinal
    quality scale with levels such as low, medium, good, and high good. We want
    to know how well the two raters actually agree with one another, beyond what
    we would expect by chance alone. Simple percent agreement can be misleading
    because some matches happen randomly. Cohen's Kappa corrects for chance
    agreement and gives us a single coefficient describing inter-rater
    reliability between counselor A and counselor B. A high kappa would reassure
    us that the rating scheme is being applied consistently across raters.
  >> VARIABLE SELECTION:
    - Rater 1: counselor_a
    - Rater 2: counselor_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's Kappa = 0.5845 (moderate agreement)   p < .001 ***
    Two counselors' assessment

>> COMMENTARY (narration):
    We measured the agreement between two counselors independently assessing the same participants with Cohen's Kappa.
    Kappa = 0.58 is in the "moderate" range -- a chance-corrected measure of agreement, i.e. the real consistency
    after removing accidental agreement. This value shows partial consistency in counselors' assessments but room for
    improvement (below 0.60 warrants attention). In clinical/counseling assessment this matters: low agreement says
    the assessment criteria should be clarified. Reporting inter-rater reliability is the quality-assurance step of
    psychological assessment.

====================================================================================

#26  Cochran-Mantel-Haenszel Test
    file: 26_cmh.xlsx
  >> SCENARIO (narration):
    Suppose we are evaluating whether an intervention improves the chance of a
    successful outcome, but we know that three different communities, A, B, and
    C, may have very different baseline success rates. In this dataset of 360
    participants we have the group assignment, control or intervention, the
    outcome, successful yes or no, and the community as a stratifying variable.
    We want to test the intervention effect while controlling for community, so
    that differences between communities do not confound our conclusion. The
    Cochran-Mantel-Haenszel Test pools the two-by-two tables across the
    community strata and tests for an overall association between group and
    success, adjusting for community.
  >> VARIABLE SELECTION:
    - Row variable: group
    - Column variable: successful
    - Stratum (layer) variable: community

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi-square(1) = 19.67   p < .001 ***   Common OR (MH) = 2.77  (95% CI: 1.76 - 4.37)
    Breslow-Day chi-square(2) = 0.04, p = 0.98 (homogeneous OR)   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between intervention group and success while controlling for communities (strata)
    with the CMH test. Even holding community constant the association is significant: common odds ratio 2.77
    (p < .001) -- stripped of the community confounder, the intervention group's odds of success are 2.8 times the
    control's. The Breslow-Day test is non-significant (p = 0.98), meaning the association is consistent across all
    communities (homogeneous OR). CMH is the classic way -- very valuable in social research -- to control for a
    third variable (community, region) by stratification.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear.xlsx
  >> SCENARIO (narration):
    Consider a study of 90 participants where we recorded three categorical
    variables at once: sex (female or male), perceived sound level (low or
    high), and whether the person is an institutional reader (yes or no). Rather
    than testing each pair of variables separately, we want to understand the
    full pattern of associations and interactions among all three at the same
    time. Log-Linear Analysis models the cell frequencies of this multi-way
    contingency table and lets us identify which main effects and interaction
    terms are needed to explain the observed counts. It is the natural tool when
    we have three or more categorical variables and care about their joint
    structure.
  >> VARIABLE SELECTION:
    - Factor 1: sex
    - Factor 2: sound
    - Factor 3: institution_reader

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Best model AIC = 46.72   Pearson chi-square = 0.0 (model fits the data perfectly)
    Factors: sex x socio-economic level x institution preference

>> COMMENTARY (narration):
    We analysed the contingency structure formed jointly by three categorical variables -- sex, socio-economic level
    and institution preference -- with a log-linear model. The selected model fits perfectly (Pearson chi-square = 0),
    with AIC = 46.72 as the most parsimonious model. Log-linear analysis lets us untangle interactions among more
    than two categorical variables (which are dependent, which interaction is significant) -- it is like ANOVA for
    categorical data. In sociology it is powerful for summarising multi-way demographic/preference relationships.

====================================================================================

#28  Crosstabulation
    file: 28_cross.xlsx
  >> SCENARIO (narration):
    Suppose we surveyed 350 respondents about their political orientation,
    classified as conservative, liberal, uncertain, or soloist, and we also
    recorded their age group as 18-30, 31-50, or 51+. Before running any formal
    test, we simply want to see how these two variables cross-tabulate, that is,
    how political orientation breaks down within each age band, with counts and
    percentages. Crosstabulation produces that contingency table along with row
    and column percentages, giving us a clear descriptive picture of how
    political attitudes vary across generations. It is the starting point for
    exploring the relationship between age and political orientation.
  >> VARIABLE SELECTION:
    - Row variable: age
    - Column variable: political

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(6) = 9.29   p = 0.158   Cramer V = 0.12 (Weak)
    NO significant association between age group and political attitude.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We examined the association between age group and political attitude with a crosstab and chi-square. This time the
    result is non-significant (chi-square(6) = 9.29, p = 0.158, Cramer V = 0.12): there is no significant association
    between age and political attitude. Non-significant results are valuable too -- always expecting an association is
    a mistake. Here age is not a factor that distinguishes political attitude; it is probably driven by other factors
    (education, class, environment). The crosstab is the most basic and readable way to see the co-occurrence of two
    categorical variables; Cramer V gives the true strength.

====================================================================================

#29  Multiple Response (MR) Frequency
    file: 29_mr_frequency.xlsx
  >> SCENARIO (narration):
    Imagine we asked 250 participants to check off all the leisure activities
    they take part in, so each person could select several options at once from
    a long list of activities. Because respondents can pick more than one, this
    is multiple-response data, and we cannot simply tally a single column. We
    want to know which activities are most popular overall, both as a percentage
    of all responses given and as a percentage of the people who answered.
    Multiple Response Frequency handles exactly this kind of pick-any data,
    reporting how often each activity was chosen across the whole sample.
  >> VARIABLE SELECTION:
    - Multiple response set (items): activities

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)
    Social activities (sports, volunteering, arts, associations...) selected as multi-choice

>> COMMENTARY (narration):
    Participants were asked which social activities they take part in -- a "multiple response" question where several
    options can be ticked. All 250 participants selected at least one activity, so the response rate is 100%.
    Multiple-response frequency analysis shows how many times and by what percentage of participants each activity was
    chosen; the totals exceed 100% because everyone can tick several boxes. It is indispensable in sociological
    studies measuring social participation, leisure and civil-society activities.

====================================================================================

#30  Multiple Response Crosstab
    file: 30_mr_categorical.xlsx
  >> SCENARIO (narration):
    Suppose we want to see whether men and women differ in the leisure
    activities they take part in, where each respondent could choose many
    activities at once. In this dataset of 220 participants we have sex, female
    or male, and a multiple-response set of activities that each person
    selected. The question is how the pattern of chosen activities differs
    between the two sex groups. Multiple Response Crosstab cross-tabulates a
    multiple-response set against a single categorical variable, letting us
    compare the activity profiles of women and men side by side.
  >> VARIABLE SELECTION:
    - Multiple response set (items): activities
    - Column (group) variable: sex

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: sex
    Social activities crossed by sex

>> COMMENTARY (narration):
    We cross-tabulated social activities (multiple-choice) by the participant's sex. With data from 220 participants
    we can see how each activity is distributed across sexes. The multiple-response crosstab answers "which group
    ticks which options more" -- e.g. women lead in one activity, men in another. This reveals the gender patterns in
    social participation and guides inclusive social-policy design.

====================================================================================

#31  MR x MR Crosstab
    file: 31_mr_mr.xlsx
  >> SCENARIO (narration):
    Imagine a career-guidance study where counselors want to understand how
    young people's personal interests map onto the careers they aspire to. In
    this dataset 180 participants could tick as many options as applied for both
    their interests (for example art, literature, science, sport) and their
    target careers (such as engineer, counselor, doctor, artist), so each person
    may report several of each. We want to see which interest-by-career
    combinations co-occur most often. A multiple-response by multiple-response
    crosstab is the right tool here because both variables are multiple-response
    sets, and an ordinary crosstab would force us to pick just one answer per
    person and throw away the rest.
  >> VARIABLE SELECTION:
    - First multiple-response set (rows): interest
    - Second multiple-response set (columns): career

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180
    Interest areas x career targets (artist, counselor, doctor, engineer...) cross-tabulated

>> COMMENTARY (narration):
    This is the most advanced multiple-response analysis: we cross two separate multi-choice questions against each
    other -- the participant's interest areas and career targets. Among 180 participants we can see which interests
    pair with which target careers: e.g. socially-interested participants target counseling/social work,
    analytically-interested ones target engineering/medicine. This table reveals interest-career matches, providing
    powerful data for career guidance and the sociology of occupations.

====================================================================================

#32  Cochran's Q Test
    file: 32_cochran_q.xlsx
  >> SCENARIO (narration):
    Suppose a guidance center piloted three different programs with the same
    group of 150 participants and asked, after each one, simply whether they
    liked it or not. Because the same people rated all three programs, the yes
    and no answers are repeated measures on a binary outcome. We want to know
    whether the proportion of people who liked the program differs across the
    three versions. Cochran's Q test is appropriate because the outcome is
    dichotomous (liked: yes or no) and we are comparing three or more related
    conditions measured on the same individuals.
  >> VARIABLE SELECTION:
    - Subject id: participant_id
    - Condition / treatment: program
    - Binary outcome (success = yes): liked

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000   (NO difference in liking proportion across programs)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the same participants' "liking" (yes/no) proportions for different programs with Cochran's Q test --
    the multi-occasion counterpart of repeated binary measures. The result is non-significant (Q = 0, p = 1.0): there
    is no difference in liking proportion across programs, participants like the programs equally. A non-significant
    result is valuable too -- it says there is no distinction in program preference in this data. Cochran's Q is the
    right way to test change in repeated binary (yes/no) measurements on the same people; finding no expected
    difference is a genuine result.

====================================================================================

#33  Correlation Analysis
    file: 33_correlation_5variable.xlsx
  >> SCENARIO (narration):
    Consider a well-being study in which 180 participants completed scales for
    motivation, weekly activity hours, psychological well-being, anxiety, and
    life satisfaction. Before building any model we want a clear picture of how
    these constructs hang together: does motivation rise with well-being, does
    anxiety move in the opposite direction, and so on. A correlation analysis is
    the natural starting point because all five variables are continuous and we
    are interested in the strength and direction of their pairwise linear
    relationships rather than predicting one specific outcome.
  >> VARIABLE SELECTION:
    - Variables: motivation
    - Variables: activity_hour
    - Variables: wellbeing
    - Variables: anxiety
    - Variables: satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    5 variables (motivation, activity hours, well-being, anxiety, satisfaction)
    NO strong relationship (|r| >= 0.75) -- relationships are moderate/weak

>> COMMENTARY (narration):
    We computed all pairwise correlations among five key psychosocial variables. No pair exceeds the |r| >= 0.75
    threshold -- so the relationships are moderate or weak. This is typical in psychology: motivation, well-being,
    anxiety are interrelated but none alone determines another; psychological outcomes are multi-factorial. The
    absence of strong relationships also means collinearity risk is low when building a regression. The correlation
    matrix is the first step before modelling to see the structure among variables.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman Analysis
    file: 34_bland_altman.xlsx
  >> SCENARIO (narration):
    Picture a measurement study where 60 participants each took two different
    cognitive tests, A and B, that are meant to capture the same underlying
    ability. The question is not whether the two scores are correlated, but
    whether they actually agree closely enough to be used interchangeably. A
    Bland-Altman analysis is the right choice because it plots the difference
    between the two methods against their average and quantifies the bias and
    the limits of agreement, which is exactly what we need when comparing two
    measurement instruments on the same people.
  >> VARIABLE SELECTION:
    - Method 1 (M1): cognitive_test_a
    - Method 2 (M2): cognitive_test_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cognitive test A mean = 103.57   |   test B mean = 101.65   |   Bias ~ +1.9 (methods agree largely)

>> COMMENTARY (narration):
    We applied two cognitive/psychological tests to the same participants and compared them with Bland-Altman. This
    method does not look at correlation (high correlation does not guarantee agreement); it measures the difference
    between the two measurements. The means are very close (103.6 vs 101.7), the bias about 2 points -- the two tests
    give almost the same result systematically and are practically interchangeable. Bland-Altman is the standard way
    to assess the agreement of two measurement tools (two tests, two raters, old/new form); it is ideal for showing
    test equivalence in psychometrics.

====================================================================================

#35  Effect Size
    file: 35_effect_size.xlsx
  >> SCENARIO (narration):
    Suppose we ran a small learning experiment where 60 participants were
    assigned to study from either a book or a video, and we recorded their
    achievement score afterwards. Beyond simply asking whether the two groups
    differ, we want to know how large that difference really is in standardized
    terms, so the result is meaningful and comparable to other studies. An
    effect size analysis is appropriate because we have a continuous outcome
    (score) and a two-level grouping variable (material: book or video), and
    effect sizes such as Cohen's d express the magnitude of the group difference
    rather than just its statistical significance.
  >> VARIABLE SELECTION:
    - Group (factor): material
    - Dependent variable: score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.93  (large effect)
    Score difference between two material groups

>> COMMENTARY (narration):
    We measured the score difference between two intervention/material groups not just as "is it significant" but
    "how large is it" -- as an effect size. Cohen's d = -0.93 is a large effect; the two groups' distributions clearly
    separate. P-values inflate with sample size, but effect size is sample-independent and gives the practical
    importance of the real difference. In psychology, "how much difference does this intervention make" is critical
    for clinical and applied decisions; a large effect of d = 0.93 says the intervention provides a strong benefit.
    Reporting effect size in publications is standard.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca.xlsx
  >> SCENARIO (narration):
    Imagine a study of 150 participants that measured two distinct domains: a
    set of cognitive ability scores (math, science, language, verbal, logic) and
    a set of psychological attitudes (motivation, attitude, satisfaction,
    anxiety). Rather than running many separate correlations, we want to know
    whether the cognitive set as a whole shares a common dimension with the
    psychological set as a whole. Canonical correlation analysis is the
    appropriate technique because it looks for linear combinations of the two
    multivariate sets that are maximally correlated with each other, summarizing
    the overall relationship between cognition and attitudes.
  >> VARIABLE SELECTION:
    - Set X: math
    - Set X: science
    - Set X: language
    - Set X: verbal
    - Set X: logic
    - Set Y: motivation
    - Set Y: attitude
    - Set Y: satisfaction
    - Set Y: anxiety

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.985, chi-square(20) = 1407.32, p < .001   |   CC2: r = 0.965, p < .001
    X set: cognitive (math/science/language/verbal/logic)  |  Y set: affective (motivation/attitude/satisfaction/anxiety)

>> COMMENTARY (narration):
    We related two multivariate sets -- cognitive skills and affective variables (motivation, attitude, satisfaction,
    anxiety) -- with canonical correlation. The first canonical axis is extremely strong (r = 0.99, p < .001), the
    second also significant (r = 0.97): the cognitive and affective domains are very tightly linked. So cognitive
    skills and motivation/attitude move together. CCA is the most direct answer to "how do two blocks of variables
    co-vary"; it is a powerful method for studying cognitive-affective interplay in psychology.

====================================================================================

#37  Correspondence Analysis (CA)
    file: 37_ca.xlsx
  >> SCENARIO (narration):
    Consider a workplace survey of 380 participants where we recorded each
    person's occupation (engineer, counselor, civil servant, doctor, or
    psychological practitioner) and the noise level of their environment,
    classified as low, medium, or high. We want to visualize how these two
    categorical variables associate, for instance whether certain occupations
    tend to cluster with louder or quieter settings. Correspondence analysis is
    well suited here because both variables are categorical and CA maps the row
    and column categories into a shared low-dimensional space, making the
    pattern of associations easy to interpret graphically.
  >> VARIABLE SELECTION:
    - Row variable: occupation
    - Column variable: sound

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 0.846   Number of dimensions = 2
    Socio-economic level x occupation preference contingency table

>> COMMENTARY (narration):
    We projected the contingency table of socio-economic level by occupation preference into a two-dimensional map
    with correspondence analysis. Total inertia 0.846 indicates a very strong association; the two dimensions
    visualise most of that structure. Level-occupation pairs that fall close on the map indicate that socio-economic
    group leans toward that occupation. In sociology this lets us read at a glance how socio-economic origin shapes
    occupational expectations -- the debate on social mobility and reproduction. It is the most powerful way to
    visualise categorical relationships.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus.xlsx
  >> SCENARIO (narration):
    Suppose a questionnaire administered to 180 participants contains eighteen
    items intended to tap three constructs: self-efficacy, motivation, and
    anxiety, each measured with six items. Before treating them as three scales,
    we want the data itself to tell us which items naturally group together.
    Variable clustering is the right approach because it works directly on the
    set of continuous item variables and partitions them into clusters of
    mutually related variables, helping us confirm whether the items split into
    the expected self-efficacy, motivation, and anxiety bundles.
  >> VARIABLE SELECTION:
    - Variables: self_efficacy_1
    - Variables: self_efficacy_2
    - Variables: self_efficacy_3
    - Variables: self_efficacy_4
    - Variables: self_efficacy_5
    - Variables: self_efficacy_6
    - Variables: motivation_1
    - Variables: motivation_2
    - Variables: motivation_3
    - Variables: motivation_4
    - Variables: motivation_5
    - Variables: motivation_6
    - Variables: anxiety_1
    - Variables: anxiety_2
    - Variables: anxiety_3
    - Variables: anxiety_4
    - Variables: anxiety_5
    - Variables: anxiety_6

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: self-efficacy / motivation / anxiety)
    A representative item was selected from each cluster

>> COMMENTARY (narration):
    We clustered 18 scale items by similarity into three groups. VarClus clusters items, not observations: it shows
    which items measure the same "construct". The three clusters formed exactly as expected: self-efficacy, motivation
    and anxiety items clustered separately. This confirms the scale really measures three sub-dimensions and lets us
    pick one representative per group to reduce the number of items. In psychological scale development and item
    reduction it is extremely practical for weeding out redundancy.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_wellbeing.xlsx
  >> SCENARIO (narration):
    Imagine we want to explain psychological well-being in a sample of 200
    participants using several lifestyle and background factors at once: how
    active they are, how motivated they feel, their parents' education level,
    and a withdrawal score. The goal is to estimate the unique contribution of
    each predictor while holding the others constant. Multiple linear regression
    is appropriate because the outcome (well-being) is continuous and we have
    several continuous predictors whose joint and individual effects on that
    outcome we want to quantify.
  >> VARIABLE SELECTION:
    - Dependent variable: wellbeing
    - Predictors: activity
    - Predictors: motivation
    - Predictors: parent_education
    - Predictors: withdrawal

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.616   Adjusted R2 = 0.608
    Well-being ~ activity + motivation + parent education + withdrawal

>> COMMENTARY (narration):
    We built a multiple linear regression predicting well-being from four variables -- social activity, motivation,
    parent education, withdrawal. The model is moderately strong: together they explain 62% of well-being variance
    (R2 = 0.62). This is a good explanatory rate in psychology -- well-being is multi-factorial and 62% captures a
    meaningful portion. Such a model shows which variable contributes most to well-being and sets intervention
    priorities. Multiple regression is the basic tool for explaining a continuous psychological outcome with several
    drivers.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic.xlsx
  >> SCENARIO (narration):
    Suppose an admissions study of 300 participants recorded whether each person
    attained their target institution, coded as a binary outcome, along with
    their well-being, activity level, motivation, and a noise exposure score. We
    want to model the probability of attainment and see which factors raise or
    lower the odds of success. Logistic regression is the correct method because
    the outcome (institution_attainment) is binary while the predictors are
    continuous, so we model the log-odds of success rather than a raw score.
  >> VARIABLE SELECTION:
    - Dependent variable (binary target): institution_attainment
    - Predictors: wellbeing
    - Predictors: activity
    - Predictors: motivation
    - Predictors: sound

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: Logistic Regression   Pseudo R2 = 0.171
    Institution attainment (1/0) ~ well-being + activity + motivation + socio-economic

>> COMMENTARY (narration):
    We built a logistic regression predicting a participant's probability of institution attainment from well-being,
    activity, motivation and socio-economic level. Because the outcome is binary (attained/not), logistic regression
    is the right choice. Pseudo R2 = 0.17 shows the model partly explains the outcome -- the coefficients generally
    say the odds of attainment rise as well-being and motivation increase. In the social sciences, logistic
    regression is the most common way to predict an outcome (attainment, success) from risk/protective factors; it
    gives each factor's effect on the odds.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Poisson / Negative Binomial Regression
    file: 41_poisson.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand what drives the number of conduct or
    disciplinary incidents a person accumulates. In this dataset of 180
    participants we recorded each person's age, their motivation score, and
    conduct_count, the count of recorded conduct incidents. Because the outcome
    is a count of events rather than a continuous score, ordinary linear
    regression is not appropriate. We use Poisson or Negative Binomial
    regression to model how age and motivation relate to the expected incident
    count, and the negative binomial variant lets us handle overdispersion if
    the counts are more variable than a pure Poisson process would predict.
  >> VARIABLE SELECTION:
    - Dependent variable (count): conduct_count
    - Predictors: age
    - Predictors: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 180.91   Deviance = 113.68
    Conduct count ~ age + motivation   (Poisson)

>> COMMENTARY (narration):
    We built a Poisson regression predicting participants' conduct-incident count -- a count variable -- from age and
    motivation. Count data (0,1,2,...) are not normally distributed, so we use Poisson rather than linear regression.
    The model typically shows conduct incidents falling as motivation rises. Deviance and AIC assess model fit. In
    the social sciences, count regression is the correct way to relate count outcomes (conduct incidents, absences,
    applications) to predictors; it is far more appropriate than a linear model.

====================================================================================

#42  Multinomial Logistic Regression
    file: 42_multinomial.xlsx
  >> SCENARIO (narration):
    Imagine we want to know what predicts which broad occupational identity a
    young person ends up choosing. Here 250 participants reported their math and
    art ability scores, and their chosen occupation falls into one of three
    unordered categories: doctor, engineer, or artist. Since the outcome is a
    categorical variable with more than two levels that have no natural
    ordering, we use multinomial logistic regression. This lets us estimate how
    math and art scores shift the relative odds of choosing each occupation
    compared to a reference category.
  >> VARIABLE SELECTION:
    - Dependent variable (nominal): occupation
    - Predictors: math
    - Predictors: art

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 252.96   Reference class: 'artist'
    Occupation preference ~ math + art score

>> COMMENTARY (narration):
    We modelled participants' occupation preference (several categories) from math and art scores with multinomial
    logistic regression. Because there are more than two unordered categories, this method fits: 'artist' is the
    reference, and a separate equation is built for each other occupation. Coefficients read as "how many times the
    odds of becoming an engineer rather than an artist change as the math score rises". This shows how the ability
    profile shapes occupational orientation -- a valuable tool in the sociology of occupations and career guidance.

====================================================================================

#43  Ordinal Logistic Regression
    file: 43_ordinal.xlsx
  >> SCENARIO (narration):
    Suppose we are studying what shapes a person's overall motivation level,
    which was rated on an ordered scale with the categories low, medium, high,
    and high high. For 200 participants we also measured counselor_support and
    community_opportunity as continuous predictors. Because the outcome is
    ordinal, the categories have a meaningful order but the gaps between them
    are not necessarily equal, we use ordinal logistic regression. This
    proportional-odds model estimates how more counselor support and greater
    community opportunity increase the odds of being in a higher motivation
    category.
  >> VARIABLE SELECTION:
    - Dependent variable (ordinal): motivation_level
    - Predictors: counselor_support
    - Predictors: community_opportunity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 466.81
    Motivation level (low<medium<high) ~ counselor support + community opportunity

>> COMMENTARY (narration):
    We built an ordinal logistic regression predicting motivation level (low, medium, high -- ordered) from counselor
    support and community opportunity. These classes are ordered, and the ordinal model uses that order, making it
    more powerful and interpretable than multinomial. The coefficients give a participant's tendency to move to a
    higher motivation level as support and opportunity increase. For ordered categorical outcomes (motivation level,
    Likert, satisfaction tier), ordinal logistic is the right method.

====================================================================================

#44  PLS Regression
    file: 44_pls.xlsx
  >> SCENARIO (narration):
    Imagine we have a battery of twelve behavioral indicators and want to
    predict an overall attainment outcome, but the behavior items are highly
    correlated with one another. In this dataset 200 participants were measured
    on behavior_01 through behavior_12 plus a continuous attainment score. When
    predictors are numerous and strongly collinear, ordinary regression becomes
    unstable, so we use Partial Least Squares regression. PLS extracts a small
    number of latent components that capture the shared variance among the
    behaviors while maximizing their relationship with attainment.
  >> VARIABLE SELECTION:
    - Dependent variable (Y): attainment
    - Predictors (X): behavior_01
    - Predictors (X): behavior_02
    - Predictors (X): behavior_03
    - Predictors (X): behavior_04
    - Predictors (X): behavior_05
    - Predictors (X): behavior_06
    - Predictors (X): behavior_07
    - Predictors (X): behavior_08
    - Predictors (X): behavior_09
    - Predictors (X): behavior_10
    - Predictors (X): behavior_11
    - Predictors (X): behavior_12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2_train = 0.877   R2_CV (5-fold) = 0.860
    Attainment ~ 12 behavior items

>> COMMENTARY (narration):
    We predicted attainment from 12 inter-correlated behavior items with PLS (Partial Least Squares) regression. When
    behavior items are correlated, ordinary regression becomes unstable; PLS solves this by reducing the items to a
    few orthogonal components. The model is very strong both in training (R2 = 0.88) and cross-validation (R2_CV =
    0.86) -- no overfitting, good generalisation. When there are many correlated items/behaviors (observation scales,
    survey batteries), PLS keeps both predictive power and interpretability.

====================================================================================

#45  Probit Regression
    file: 45_probit.xlsx
  >> SCENARIO (narration):
    Suppose we want to know whether the amount of time someone spends on
    activities affects their probability of passing a threshold. For 180
    participants we recorded activity_hour and a binary passed indicator.
    Because the outcome is a yes or no event, we model the probability of
    passing using probit regression, which links the predictor to the outcome
    through the cumulative normal distribution. This tells us how each
    additional hour of activity raises the likelihood that a participant passes.
  >> VARIABLE SELECTION:
    - Dependent variable (binary): passed
    - Predictors: activity_hour

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 97.53   Pseudo R2 (McFadden) = 0.603
    Passing (1/0) ~ activity hours

>> COMMENTARY (narration):
    We predicted whether a participant passes from activity hours with probit regression. Probit models a binary
    outcome like logistic but assumes a normal-distribution curve (cumulative normal). Pseudo R2 = 0.60 is very
    strong: activity hours largely explain the probability of passing. Probit and logistic usually give similar
    results; the choice is often tradition and interpretive preference. In the social sciences it suits predicting
    success/passing probability from a continuous predictor.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit.xlsx
  >> SCENARIO (narration):
    Imagine we are studying how much families spend on education tuition, but
    many families report zero spending, so the outcome is censored at the lower
    bound. In this dataset of 200 families we have household income,
    child_count, and education_tuition as the dependent variable. Ordinary
    regression would be biased by the pile-up of zeros, so we use Tobit
    regression, which explicitly models the censoring and recovers the
    underlying relationship between income, number of children, and how much
    families would spend on tuition.
  >> VARIABLE SELECTION:
    - Dependent variable (censored): education_tuition
    - Predictors: income
    - Predictors: child_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 475.35   Tobit (left-censored: expenditure floored)
    Education tuition ~ income + child count

>> COMMENTARY (narration):
    We modelled families' education tuition from income and number of children. The problem: some families spend
    nothing on education (zero) or expenditure piles up below a floor value -- this is censored data, and ordinary
    regression gives biased results. Tobit regression correctly models this left-censored structure (AIC = 475.35).
    Typically expenditure rises with income and falls with more children. The statistically correct solution for
    working with censored/piled-up data (expenditure, duration) is the Tobit model.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict a person's score from their activity level and
    age, but we also want to express our uncertainty about the coefficients as
    full probability distributions rather than single point estimates. For these
    100 participants we use Bayesian linear regression, which combines prior
    beliefs with the observed data to produce posterior distributions for each
    effect. This approach is especially useful with a modest sample size because
    it gives us credible intervals and lets us quantify exactly how confident we
    are about the influence of activity and age on score.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Predictors: activity
    - Predictors: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Conjugate Normal-Inverse-Gamma prior   posterior sigma2 = 16.95
    Score ~ activity + age

>> COMMENTARY (narration):
    We used Bayesian linear regression to predict score from activity and age. The Bayesian advantage: the result is
    not a single p-value but a probability distribution (posterior) for the parameters -- we get a credible interval
    for each coefficient and a "probability the effect is positive". The posterior error variance is sigma2 = 16.95.
    The Bayesian method honestly expresses uncertainty and can incorporate prior knowledge. It is valuable in
    psychological research with small samples or where information must be carried over from previous studies.

====================================================================================

#48  Nonlinear Regression
    file: 48_nonlinear.xlsx
  >> SCENARIO (narration):
    Imagine we are tracking how children's vocabulary grows with age. For 80
    children we recorded their age and their word_count. Language acquisition
    does not increase in a straight line, it tends to follow a curved,
    saturating growth pattern, so a linear model would fit poorly. We use
    nonlinear regression to fit a curve in which word_count is modeled as a
    nonlinear function of age, capturing the accelerating then leveling shape of
    early vocabulary development.
  >> VARIABLE SELECTION:
    - Dependent variable (Y): word_count
    - Independent variable (X): age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: y = K / (1 + exp(-r·(x-x0)))  (logistic growth)   R2 = 0.966
    Language (word count) ~ age

>> COMMENTARY (narration):
    We modelled children's vocabulary growth by age. Language acquisition is not linear: it starts slow, rises fast,
    then saturates toward a ceiling (K) -- a classic S-curve (logistic growth). We modelled this with a logistic
    function and the fit is excellent (R2 = 0.97). The curve shows at which age development accelerates and when it
    approaches the ceiling. Nonlinear regression is the way to correctly model growth/learning curves in
    developmental psychology -- language acquisition, skill gain; forcing a straight line would misrepresent the
    nature of development.

====================================================================================

#49  Ridge Regression
    file: 49_ridge.xlsx
  >> SCENARIO (narration):
    Suppose we want to predict motivation from a psychological questionnaire
    with fifteen items, but the items are closely related and highly collinear.
    In this dataset 200 participants answered item_p01 through item_p15, with
    motivation as the continuous outcome. With so many correlated predictors,
    ordinary least squares inflates the coefficients and overfits, so we use
    Ridge regression. Ridge applies an L2 penalty that shrinks the coefficients
    toward zero, stabilizing the estimates while keeping every item in the
    model.
  >> VARIABLE SELECTION:
    - Dependent variable: motivation
    - Predictors: item_p01
    - Predictors: item_p02
    - Predictors: item_p03
    - Predictors: item_p04
    - Predictors: item_p05
    - Predictors: item_p06
    - Predictors: item_p07
    - Predictors: item_p08
    - Predictors: item_p09
    - Predictors: item_p10
    - Predictors: item_p11
    - Predictors: item_p12
    - Predictors: item_p13
    - Predictors: item_p14
    - Predictors: item_p15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Ridge (alpha = 1.0): R2 = 0.618   Adj. R2 = 0.586   n = 200
    Motivation ~ 15 scale items (mutually correlated)

>> COMMENTARY (narration):
    Predicting motivation from 15 highly correlated scale items, ordinary regression becomes unstable (collinearity
    inflates coefficients). Ridge regression solves this by shrinking all coefficients -- it does not zero them but
    limits their magnitude, keeping prediction stable (R2 = 0.62). Ridge is ideal when "I want to keep all items but
    there is collinearity". For naturally correlated variable sets like scale items, it prevents overfitting and
    provides stable estimates.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso.xlsx
  >> SCENARIO (narration):
    Imagine we have thirty candidate predictors and we suspect that only a
    handful of them actually matter for attainment. Here 250 participants were
    measured on x01 through x30 plus a continuous attainment outcome. When we
    want both prediction and automatic variable selection, we use Lasso
    regression. The L1 penalty shrinks unimportant coefficients exactly to zero,
    so Lasso identifies the small subset of predictors that genuinely drive
    attainment while discarding the rest.
  >> VARIABLE SELECTION:
    - Dependent variable: attainment
    - Predictors: x01
    - Predictors: x02
    - Predictors: x03
    - Predictors: x04
    - Predictors: x05
    - Predictors: x06
    - Predictors: x07
    - Predictors: x08
    - Predictors: x09
    - Predictors: x10
    - Predictors: x11
    - Predictors: x12
    - Predictors: x13
    - Predictors: x14
    - Predictors: x15
    - Predictors: x16
    - Predictors: x17
    - Predictors: x18
    - Predictors: x19
    - Predictors: x20
    - Predictors: x21
    - Predictors: x22
    - Predictors: x23
    - Predictors: x24
    - Predictors: x25
    - Predictors: x26
    - Predictors: x27
    - Predictors: x28
    - Predictors: x29
    - Predictors: x30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Lasso (alpha = 0.1): R2 = 0.726   Adj. R2 = 0.688   n = 250
    18 of 30 predictor coefficients shrunk to zero (automatic feature selection)

>> COMMENTARY (narration):
    In this data a few real predictors were mixed with many irrelevant variables. The strength of Lasso shows here:
    the penalty term drives irrelevant variables' coefficients exactly to zero -- automatic variable selection. 18 of
    30 predictors were eliminated, and the model stayed strong at R2 = 0.73. Lasso is extremely useful when you want
    to pick the true drivers from many candidates -- in high-dimensional psychometric/social data. While Ridge
    shrinks all variables, Lasso eliminates the irrelevant ones entirely.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand how parental support translates into academic
    attainment. We suspect the effect is not direct: when parents are more
    supportive, students become more motivated, and that heightened motivation
    is what actually drives their attainment. In this dataset of 200
    participants we measured parent_support, motivation, and attainment on
    continuous scales. Mediation analysis is the right tool here because it lets
    us decompose the total effect of parent_support on attainment into a direct
    path and an indirect path that runs through motivation, telling us how much
    of the relationship is explained by the mediator.
  >> VARIABLE SELECTION:
    - Independent variable (X): parent_support
    - Mediator (M): motivation
    - Dependent variable (Y): attainment

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 5.64   p < .001   95% CI (3.91 , 7.57)   Significant
    Path: parent support -> motivation -> attainment

>> COMMENTARY (narration):
    We tested whether parent support's effect on attainment passes through motivation with mediation analysis. The
    indirect effect is significant (5.64, p < .001, CI excludes zero): parent support raises motivation, and higher
    motivation raises attainment. So parent support affects attainment not directly but "through" motivation. This
    opens up one of the core questions of developmental and educational psychology -- the mechanism by which family
    influence works. Mediation analysis reveals the intermediate mechanism in "why X affects Y" and points to
    intervention targets.

====================================================================================

#52  Path Analysis
    file: 52_path.xlsx
  >> SCENARIO (narration):
    Imagine we have a theory about how students reach academic success. We think
    self_efficacy fuels motivation, motivation increases effort, and effort
    finally produces attainment, while some of these variables may also act on
    attainment directly. With 220 participants measured on self_efficacy,
    motivation, effort, and attainment, a single mediation model is too narrow.
    Path analysis is appropriate because it lets us estimate this whole web of
    directed relationships simultaneously, test the proposed causal structure,
    and see which direct and indirect pathways hold up.
  >> VARIABLE SELECTION:
    - Exogenous variable: self_efficacy
    - Endogenous/mediator: motivation
    - Endogenous/mediator: effort
    - Final outcome: attainment
    - Model paths: self_efficacy -> motivation -> effort -> attainment

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.866   RMSEA = 0.243   (poor fit -- model should be revised)
    Model: self-efficacy/motivation/effort -> attainment; self-efficacy -> motivation

>> COMMENTARY (narration):
    Path analysis is extended regression that models several cause-effect relationships at once. Here we tested that
    self-efficacy, motivation and effort affect attainment, and that self-efficacy feeds motivation. CFI = 0.87 and
    RMSEA = 0.24 -- the fit is poor, so the hypothesised model does not fully match the data; some paths may be
    missing or extra. This is informative too: it says the theoretical model should be revised. Path analysis is a
    powerful way to separate direct and indirect effects and test psychological theories; both good and poor fit guide
    us in improving the model.

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm.xlsx
  >> SCENARIO (narration):
    Suppose we tracked the same 200 participants across four assessment waves,
    labeled D1, D2, D3 and D4, recording their wellbeing each time. Because the
    repeated wellbeing scores come from the same individuals, the observations
    are not independent and ordinary regression would be misleading. A linear
    mixed model is the right choice: we treat wave as a fixed effect to capture
    how wellbeing changes over time, and we add a random intercept for each
    participant to account for the fact that some people simply start higher or
    lower than others.
  >> VARIABLE SELECTION:
    - Dependent variable: wellbeing
    - Fixed effect (factor): wave
    - Random effect / grouping (subject): participant_id

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group variance (random intercept) = 356.66   ICC = 0.943
    Well-being ~ wave  +  (1 | participant)

>> COMMENTARY (narration):
    Modelling the same participants' well-being across waves, the within-person repeated measures are not
    independent. The linear mixed model solves this by adding participant as a random effect. ICC = 0.94 is very high:
    94% of well-being variance comes from between-person differences; within-wave change is relatively small. So
    people are consistent within themselves, and the main difference is between individuals. LMM is the correct,
    indispensable method for nested/repeated-measures data (participant>wave) -- it models each person's own
    trajectory and resolves the independence violation.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation.xlsx
  >> SCENARIO (narration):
    Suppose our wellbeing survey came back with gaps: some participants skipped
    items on activity, motivation, anxiety, or parent_education, so several
    wellbeing records are incomplete. Simply deleting those cases would shrink
    our sample of 150 and could bias the results if the missingness is not
    random. Multiple imputation is the appropriate strategy here because it
    generates several plausible complete datasets using the relationships among
    wellbeing, activity, motivation, anxiety, and parent_education, then pools
    the analyses so our estimates and standard errors honestly reflect the
    uncertainty from the missing values.
  >> VARIABLE SELECTION:
    - Variable with missing values: wellbeing
    - Variable with missing values: activity
    - Variable with missing values: motivation
    - Variable with missing values: anxiety
    - Variable with missing values: parent_education

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (number of imputations) = 5
    Missing well-being values estimated from activity/motivation/anxiety/parent education

>> COMMENTARY (narration):
    Missing observations are inevitable in social data (unanswered items, lost to follow-up); deleting them both
    loses data and biases results. Multiple imputation estimates the missing well-being values from other variables,
    creating 5 separate completed datasets, runs the analysis in each and combines the results -- thereby also
    accounting for estimation uncertainty. Under the MAR (missing at random) assumption this is the modern
    missing-data standard. In longitudinal psychological/panel studies it is the soundest way to cope with lost
    participant data.

====================================================================================

#55  Generalized Estimating Equations (GEE)
    file: 55_gee.xlsx
  >> SCENARIO (narration):
    Imagine an intervention study where 300 participants were followed across
    repeated visits and we recorded their motivation at each one, comparing an
    intervention group against a control group. Since each person contributes
    several correlated motivation measurements, we cannot pretend the records
    are independent. Generalized estimating equations are well suited here: they
    let us model the population-average effect of group and visit on motivation
    while specifying a working correlation structure to handle the within-
    participant dependence across visits.
  >> VARIABLE SELECTION:
    - Dependent variable: motivation
    - Predictor (group): group
    - Time / repeated measure: visit
    - Subject (cluster) id: participant_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coefficient = 0.032   p = 0.042 *   QIC = 305.56
    Motivation ~ visit  (repeated measures, clustered within person)

>> COMMENTARY (narration):
    We took repeated (multi-visit) motivation measurements from the same participants -- one person's measurements are
    dependent. GEE estimates the relationship at the "population average" level in such clustered data, accounting for
    the correlation structure. The visit effect is borderline significant (0.032, p = 0.042): motivation rises
    slightly over time. GEE is an alternative to LMM: rather than modelling random effects, it corrects the
    correlation among observations and gives the marginal (average) effect. It is a common, powerful method for
    longitudinal/repeated psychological measurement data.

====================================================================================

#56  Generalized Linear Mixed Model (GLMM)
    file: 56_glmm.xlsx
  >> SCENARIO (narration):
    Suppose we are studying behavioral conduct in 160 participants who were
    observed over several months, where each conduct outcome is a count rather
    than a continuous score, and we want to know whether an intervention reduces
    it. Because the monthly counts are repeated within the same individuals and
    the outcome is a count, neither a plain mixed model nor a plain Poisson
    model is enough. A generalized linear mixed model fits the situation: it
    models conduct with an appropriate count distribution, includes intervention
    and month as fixed effects, and adds a random intercept per participant to
    absorb individual differences in baseline behavior.
  >> VARIABLE SELECTION:
    - Dependent variable (count): conduct
    - Fixed predictor: intervention
    - Time covariate: month
    - Random effect / grouping (subject): participant_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    intervention: coef = -0.206, p = 0.077 (borderline non-significant)
    Poisson GLMM: conduct count ~ month + intervention  +  (1 | participant)

>> COMMENTARY (narration):
    We modelled conduct-incident counts (Poisson) with both month/intervention fixed effects and a participant random
    effect -- a GLMM: mixed model + count distribution. The intervention effect is borderline non-significant (coef
    -0.21, p = 0.077): there is a negative trend (the intervention tends to reduce conduct incidents) but it does not
    cross the significance threshold. A non-significant/borderline result is information too -- a larger sample could
    clarify it. GLMM combines the strengths of LMM (which assumes normality) and GLM (which ignores clustering) for
    "repeated/nested count or binary data". It is the correct framework for monitoring behavioral interventions.

====================================================================================

#57  Regularized Regression (Elastic Net)
    file: 57_elasticnet.xlsx
  >> SCENARIO (narration):
    Suppose we collected 40 psychological indicators, x01 through x40, and want
    to predict a composite index score, but many of these predictors are
    correlated and we are not sure which truly matter. With 220 participants and
    40 candidate predictors, ordinary regression risks overfitting and unstable
    coefficients. Regularized regression with an elastic net is ideal here
    because it blends ridge and lasso penalties to shrink coefficients, handle
    the correlated predictors gracefully, and perform variable selection so we
    end up with a parsimonious, stable model.
  >> VARIABLE SELECTION:
    - Target variable: index
    - Predictors: x01, x02, x03, x04, x05, x06, x07, x08, x09, x10, x11, x12, x13, x14, x15, x16, x17, x18, x19, x20, x21, x22, x23, x24, x25, x26, x27, x28, x29, x30, x31, x32, x33, x34, x35, x36, x37, x38, x39, x40

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.075   R2 (train) = 0.604
    Index ~ 40 predictors

>> COMMENTARY (narration):
    Elastic Net is a hybrid of Ridge and Lasso: it both shrinks coefficients (like Ridge) and zeros out the
    irrelevant ones (like Lasso). Here we predicted a psychosocial index from 40 predictors; Elastic Net keeps
    correlated variable groups together while eliminating the irrelevant ones. The model gives R2 = 0.60. When there
    are many collinear predictors -- item batteries, behavior scales, demographics -- Elastic Net is a balanced choice
    offering both selection and stability. It softens Lasso's "pick one at random from a correlated group" problem.

====================================================================================

#58  Robust Regression
    file: 58_robust.xlsx
  >> SCENARIO (narration):
    Imagine we want to see how weekly activity relates to a psychological score
    across 100 participants, but a few respondents reported extreme, outlying
    values that could distort an ordinary least squares line. Rather than
    deleting those cases by hand, we want an estimator that is less sensitive to
    them. Robust regression is the appropriate method here: it down-weights the
    influence of outliers so the fitted relationship between activity and score
    reflects the bulk of the data rather than being dragged around by a handful
    of unusual points.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Predictor: activity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust intercept = 51.24   vs   OLS intercept = 49.14   (outliers distorted OLS)
    Score ~ activity

>> COMMENTARY (narration):
    This data had a few outliers. Ordinary regression (OLS) takes outliers fully into account and is pulled toward
    them: the intercept is 49.14 in OLS but 51.24 in the robust model; the marked gap means outliers distorted OLS.
    Robust regression (Huber-T) gives less weight to outlying observations to preserve the true trend. In
    psychological data, where outlier participant values are inevitable, robust regression gives a more reliable slope
    estimate than OLS.

====================================================================================

#59  Quantile Regression
    file: 59_quantile.xlsx
  >> SCENARIO (narration):
    Suppose we are interested in how activity and motivation relate to
    wellbeing, but we suspect their effect is not the same across the whole
    distribution: predictors might matter much more for people at the low end of
    wellbeing than for those who are already thriving. With 200 participants,
    focusing only on the mean would hide this. Quantile regression is the right
    approach because it lets us model wellbeing at several quantiles, revealing
    whether activity and motivation have different effects for low-, median-,
    and high-wellbeing individuals.
  >> VARIABLE SELECTION:
    - Dependent variable: wellbeing
    - Predictor: activity
    - Predictor: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    q = 0.10 slope = 0.221   |   q = 0.50 (median)   |   q = 0.90  (varying effect)
    Well-being ~ activity + motivation

>> COMMENTARY (narration):
    Ordinary regression models only the mean; but activity's effect on well-being may differ for low- and high-
    well-being people. Quantile regression models the lower (q=0.10, low well-being), middle (q=0.50) and upper
    (q=0.90, high well-being) quantiles separately. A changing slope across quantiles answers "does activity help
    everyone equally" -- for instance activity may help those with the lowest well-being more. In psychology, when
    the different parts of the distribution (most vulnerable/best-off) matter rather than the average, quantile
    regression reveals what mean-based models miss.

====================================================================================

#60  ROC Curve Analysis
    file: 60_roc.xlsx
  >> SCENARIO (narration):
    Imagine we built a composite psychological screening score and want to know
    how well it distinguishes participants who ultimately succeeded from those
    who did not. Across 250 participants we have the continuous composite score
    and a binary successful outcome. ROC curve analysis is exactly the tool for
    this: by sweeping across all possible cut-off values it shows the trade-off
    between sensitivity and specificity, gives us the area under the curve as an
    overall measure of discrimination, and helps us choose an optimal threshold
    for classifying success.
  >> VARIABLE SELECTION:
    - Test / predictor variable (continuous): composite
    - Outcome / class variable (binary): successful

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.921 (excellent discriminating power)   Youden optimum threshold: 0.71 -> Sensitivity 0.79
    n = 250 (positive 150, negative 100)   success classifier

>> COMMENTARY (narration):
    We measured how well a composite score predicts success with the ROC curve. AUC = 0.92 is "excellent": the score
    largely separates successful from unsuccessful participants correctly. The Youden index gives the best cut-off
    (threshold), optimising sensitivity and specificity; here sensitivity is 0.79 at a 0.71 threshold. ROC/AUC is the
    standard for reporting a diagnostic/classification model's discriminating power; in the social sciences it is
    ideal for evaluating risk-screening and prediction tools.

====================================================================================

#61  True Skill Statistic (TSS)
    file: 61_tss.xlsx
  >> SCENARIO (narration):
    Imagine we built a model to flag students at risk of dropping out, and now
    we want to honestly judge how good that early-warning flag really is. In
    this dataset of 200 participants we have a continuous model output called
    prediction, alongside two indicators measured during the year, wellbeing and
    withdrawal, and a binary outcome called pass that records whether the
    student actually succeeded. The True Skill Statistic is the right tool here
    because, unlike raw accuracy, it corrects for how often the positive class
    occurs and tells us how much better our prediction does than chance at
    separating those who pass from those who do not. This matters in psychology
    and education research, where the outcome is often imbalanced and a
    flattering accuracy number can hide a useless classifier.
  >> VARIABLE SELECTION:
    - Observed (actual outcome): pass
    - Predicted score / probability: prediction

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.338 (acceptable)   Sensitivity 0.74   Specificity 0.60   n = 200
    Passing/success prediction

>> COMMENTARY (narration):
    TSS (True Skill Statistic) measures how much better than chance a binary classification model is: sensitivity +
    specificity - 1. The model predicting success has TSS = 0.34 -- "acceptable": sensitivity 0.74 (good at catching
    successes), specificity 0.60 (moderate at separating failures). TSS's advantage is insensitivity to prevalence
    (how common success is), so it is reliable for imbalanced classes. It summarises the true skill of success/risk
    prediction models in the social sciences, getting past the bias that accuracy hides.

====================================================================================

#62  Confusion Matrix Metrics
    file: 62_complexity.xlsx
  >> SCENARIO (narration):
    Suppose a counseling service deployed an automated screening tool and we
    need to report exactly where it succeeds and where it fails. Here we have
    200 cases, each with an actual_label showing the true status, a
    prediction_label produced by the tool, and a continuous prediction_score
    behind that decision. Confusion Matrix Metrics let us go beyond a single
    accuracy figure and break performance down into sensitivity, specificity,
    precision and F1, so we can see whether the tool misses true cases or raises
    too many false alarms. That distinction is crucial in mental-health
    screening, where the cost of a missed case is very different from the cost
    of a false positive.
  >> VARIABLE SELECTION:
    - Actual class: actual_label
    - Predicted class: prediction_label
    - Predicted score / probability: prediction_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.93   F1 = 0.918
    Class prediction -- actual vs predicted label

>> COMMENTARY (narration):
    We broke down a classifier's predictions with a confusion matrix. With accuracy 93% and F1 = 0.92 the model is
    very successful -- both the overall correct rate is high and the precision/recall balance (F1) is good. The
    confusion matrix reveals what the "accuracy" number hides -- where each class is misclassified, which classes are
    confused. The F1 score is more informative than accuracy especially with imbalanced classes. In the social
    sciences it is the indispensable tool for evaluating the real performance of automated classification models.

====================================================================================

#63  Random Forest
    file: 63_rf.xlsx
  >> SCENARIO (narration):
    Let's say we want to predict a person's occupation from their cognitive and
    psychological profile and, just as importantly, understand which traits
    matter most. In this dataset of 350 participants we have five numeric
    predictors, math, language, wellbeing, verbal and age, and the categorical
    target occupation, which takes the values doctor, engineer, counselor and
    artist. A Random Forest is well suited here because the relationships are
    likely nonlinear and interacting, it handles a multi-class outcome
    naturally, and its variable-importance ranking tells us which abilities
    drive occupational membership without us having to specify a model form in
    advance.
  >> VARIABLE SELECTION:
    - Target (class): occupation
    - Predictors: math
    - Predictors: language
    - Predictors: wellbeing
    - Predictors: verbal
    - Predictors: age

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.914   (occupation preference classification)
    Predicted from 5 variables (math, language, well-being, verbal, age)

>> COMMENTARY (narration):
    We predicted participants' occupation preference from five variables (math, language, well-being, verbal, age)
    with Random Forest -- the vote of hundreds of decision trees. Accuracy is 91%: the variables largely classify
    occupation preference correctly. One of Random Forest's most valuable outputs is the "variable importance
    ranking", telling which variable matters most for the choice. It is a powerful, easy-to-build machine-learning
    method for nonlinear, interacting relationships; it is valuable for predicting occupational/career orientation.

====================================================================================

#64  Support Vector Machine (SVM)
    file: 64_svm.xlsx
  >> SCENARIO (narration):
    Imagine we want to classify whether a student passes an institutional
    benchmark based on their psychological and performance measures. Across 250
    participants we have four numeric predictors, wellbeing, math, trial and
    activity, and a binary target institution_passed. A Support Vector Machine
    is a strong choice here because it finds the boundary that maximally
    separates the two groups even when the classes overlap and the decision
    surface is not linear, and with a kernel it can capture curved relationships
    between traits and the pass outcome that a simple logistic model would miss.
  >> VARIABLE SELECTION:
    - Target (class): institution_passed
    - Predictors: wellbeing
    - Predictors: math
    - Predictors: trial
    - Predictors: activity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (institution attainment yes/no)
    Predicted from 4 variables (well-being, math, trial, activity)

>> COMMENTARY (narration):
    We found the best boundary separating participants who attained an institution from those who did not with SVM.
    SVM seeks the maximum-margin separating surface between two classes and, via the kernel trick, can model nonlinear
    boundaries too. Accuracy is 100% -- the variables separate the two groups perfectly (such high accuracy should be
    kept in mind regarding overfitting). SVM is especially powerful in small-to-medium, high-dimensional
    classification problems and is robust thanks to margin maximisation. It is a strong alternative to Random Forest.

====================================================================================

#65  Gradient Boosting
    file: 65_gradient_boosting.xlsx
  >> SCENARIO (narration):
    Suppose an advising office wants an accurate, tunable model that predicts
    whether a student is at low or medium risk of attrition so it can target
    support. In this dataset of 280 participants we have four numeric
    predictors, wellbeing, withdrawal, sound and parent_joint, and the
    categorical target attrition_risk with levels medium and low. Gradient
    Boosting fits here because it builds an ensemble of trees sequentially, each
    correcting the errors of the last, which usually squeezes out higher
    predictive accuracy than a single tree and still gives us feature-importance
    information about which factors push risk up.
  >> VARIABLE SELECTION:
    - Target (class): attrition_risk
    - Predictors: wellbeing
    - Predictors: withdrawal
    - Predictors: sound
    - Predictors: parent_joint

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.982   (attrition risk classification)
    well-being + withdrawal + socio-economic + parent togetherness

>> COMMENTARY (narration):
    We predicted attrition risk from psychosocial variables with Gradient Boosting. Boosting adds trees sequentially,
    each new tree correcting the previous ones' errors; accuracy is very high at 98.2%. This is critical for
    early-warning systems: high-attrition-risk people can be identified in advance and intervened with. While Random
    Forest builds trees in parallel, Boosting builds them sequentially (chasing the error); it usually gives higher
    accuracy but needs careful tuning. It is a valuable tool in social services and program-continuity management.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans.xlsx
  >> SCENARIO (narration):
    Let's say we suspect that participants naturally fall into distinct profiles
    based on their wellbeing and engagement, but we have no labels to tell us
    so. Here we have 240 participants described by four numeric measures,
    wellbeing, activity, math and trial. K-Means Clustering is appropriate
    because it is an unsupervised method that partitions people into a chosen
    number of compact groups by similarity on these traits, letting us discover,
    for example, a thriving high-engagement profile versus a low-wellbeing
    disengaged profile without imposing any prior grouping.
  >> VARIABLE SELECTION:
    - Clustering variables: wellbeing
    - Clustering variables: activity
    - Clustering variables: math
    - Clustering variables: trial

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.459 (moderate cluster structure)
    Participant typology from 4 variables (well-being, activity, math, trial)

>> COMMENTARY (narration):
    We split participants into four natural groups (a typology) by their psychosocial features with K-Means.
    Silhouette = 0.46 is moderate -- the clusters separate but not sharply, which is expected in psychological data
    (features are a continuous gradient). The four groups likely represent distinct participant profiles: e.g. high
    well-being-active, low well-being-withdrawn, and intermediate types. K-Means finds natural groups in unlabeled
    data; it is practical for participant profiling, designing targeted interventions and identifying risk groups.
    The silhouette score shows the number of clusters is reasonable.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic.xlsx
  >> SCENARIO (narration):
    Imagine we have a small, carefully observed sample and we want to see how
    individuals group together based on a rich behavioral profile, while also
    knowing their stated learning style. In this dataset of 40 participants we
    have ten numeric behavior indicators, behavior_01 through behavior_10, plus
    a categorical style variable with the levels visual, auditory, kinesthetic,
    enrolled and mixed. Hierarchical Clustering suits this situation because the
    sample is small, we do not want to fix the number of clusters in advance,
    and the dendrogram reveals the nested structure of similarity, after which
    we can check whether the emergent clusters align with the recorded learning
    styles.
  >> VARIABLE SELECTION:
    - Clustering variables: behavior_01
    - Clustering variables: behavior_02
    - Clustering variables: behavior_03
    - Clustering variables: behavior_04
    - Clustering variables: behavior_05
    - Clustering variables: behavior_06
    - Clustering variables: behavior_07
    - Clustering variables: behavior_08
    - Clustering variables: behavior_09
    - Clustering variables: behavior_10

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 5   Silhouette = 0.123 (weak/overlapping structure)
    5 styles from 10 behavior items

>> COMMENTARY (narration):
    We grouped participants by behavior profile with hierarchical clustering -- the result is a dendrogram (tree).
    Splitting into five clusters gives a silhouette of only 0.12: the clusters overlap, there is no sharp separation.
    This is common in psychology -- behavior styles are often a continuous spectrum, not sharp categories. A low
    silhouette means "natural grouping in the data is weak", which is informative (perhaps fewer than 5 styles exist).
    Hierarchical clustering is more explanatory than K-Means for discovering the nested structure of groups and the
    right number of clusters.

====================================================================================

#68  DBSCAN Clustering
    file: 68_dbscan.xlsx
  >> SCENARIO (narration):
    Suppose we are studying where people live across a city and want to find
    dense residential clusters while flagging isolated households as outliers.
    Here we have 140 communities, each with a latitude and longitude, lat and
    lon. DBSCAN is the right method because it groups points by spatial density
    rather than forcing every point into a cluster, so it can find irregularly
    shaped neighborhoods and explicitly label sparse, scattered locations as
    noise, which is exactly what we want when mapping social settlement
    patterns.
  >> VARIABLE SELECTION:
    - Coordinates (X / longitude): lon
    - Coordinates (Y / latitude): lat

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.916   (community locations)
    Communities clustered by geographic density

>> COMMENTARY (narration):
    We clustered communities' geographic locations by density with DBSCAN. Unlike K-Means: we do not give the number
    of clusters in advance, and it flags low-density isolated points as "noise" (outliers). Four dense community
    groups were found (silhouette 0.92, very strong). This shows communities cluster in particular regions/
    neighbourhoods. DBSCAN is more appropriate than K-Means for spatial data with irregularly shaped clusters and
    outliers (community locations, household addresses); it is used to analyse social geography and community
    structure.

====================================================================================

#69  Principal Component Analysis (PCA)
    file: 69_pca.xlsx
  >> SCENARIO (narration):
    Let's say we measured six cognitive abilities and suspect they really
    reflect a smaller number of underlying dimensions. In this dataset of 180
    participants we have math, science, language, verbal, logic and art.
    Principal Component Analysis is ideal here because it reduces these
    correlated measures into a few orthogonal components that capture most of
    the variance, helping us see, for instance, whether a general ability
    dimension and a separate verbal-versus-quantitative contrast emerge, and
    giving us compact scores to use in later analyses.
  >> VARIABLE SELECTION:
    - Variables: math
    - Variables: science
    - Variables: language
    - Variables: verbal
    - Variables: logic
    - Variables: art

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 = 88.3%   (the first component alone explains most of the variance)
    6 skill domains (math, science, language, verbal, logic, art)

>> COMMENTARY (narration):
    We reduced six skill domains to a few summary axes with PCA. The first component alone explains 88.3% of the
    variance -- strikingly high. The meaning: domains are very highly correlated, all measuring a single "general
    ability" dimension (g-factor-like). So a person good at one domain is generally good at the others. PCA is the
    basic tool for summarising multivariate data, reducing collinearity and revealing hidden dimensions (like general
    ability). One component being this dominant is a concrete reflection of the "general intelligence" (g) debate in
    psychology.

====================================================================================

#70  t-SNE
    file: 70_tsne.xlsx
  >> SCENARIO (narration):
    Imagine we have a twelve-item behavioral questionnaire and we want to
    visualize whether respondents of different known types separate into
    recognizable clusters. Here we have 125 participants, each answering item_01
    through item_12, along with a categorical type label taking the values A, B,
    C, D and E. t-SNE is well suited because it is a nonlinear dimensionality-
    reduction technique designed for visualization, projecting the high-
    dimensional item space onto two dimensions while preserving local
    neighborhood structure, so we can color the points by type and visually
    inspect whether the response patterns cluster as expected.
  >> VARIABLE SELECTION:
    - Variables: item_01
    - Variables: item_02
    - Variables: item_03
    - Variables: item_04
    - Variables: item_05
    - Variables: item_06
    - Variables: item_07
    - Variables: item_08
    - Variables: item_09
    - Variables: item_10
    - Variables: item_11
    - Variables: item_12
    - Color / label (optional): type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence (fit quality) = 0.544 -> good
    12-item profile (participant types) reduced to 2D

>> COMMENTARY (narration):
    t-SNE is a nonlinear method that projects high-dimensional participant profile data (12 items) into a
    two-dimensional plot. Unlike PCA, it focuses on preserving local neighbourhoods -- similar participants are placed
    close on the map -- ideal for seeing hidden group structure. The KL divergence is 0.54, so the reduction quality
    is good; the participant types form visible clusters on the map. Important caveat: t-SNE axes and between-cluster
    distances are not interpretable, only the clustering pattern is. It is a powerful tool for visually exploring
    psychological types/profiles.

====================================================================================

#71  Multidimensional Scaling (MDS)
    file: 71_mds.xlsx
  >> SCENARIO (narration):
    Suppose we want to understand how 51 participants relate to one another
    across several psychological and academic dimensions, and we'd like a map
    that shows which people are similar and which are far apart. For each
    participant we have wellbeing, weekly activity, and three ability scores in
    math, science, and language. Rather than looking at each variable
    separately, we want to compress these differences into a two-dimensional
    picture where the distance between points reflects overall dissimilarity
    between participants. Multidimensional Scaling is ideal here because it
    takes the full set of continuous profiles, computes pairwise distances, and
    projects them onto a low-dimensional space we can actually visualize and
    interpret.
  >> VARIABLE SELECTION:
    - Coordinates (input variables): wellbeing
    - Coordinates (input variables): activity
    - Coordinates (input variables): math
    - Coordinates (input variables): science
    - Coordinates (input variables): language

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.052 (acceptable/good fit)
    Participant similarity map (well-being, activity, 3 skills)

>> COMMENTARY (narration):
    MDS maps the similarity among participants into two dimensions -- participants with similar profiles placed near,
    dissimilar ones far. Stress = 0.052 is a "good" fit; this means the participant-similarity structure fits two
    dimensions successfully (low stress = reliable map). Participants close on the map have similar psychosocial
    profiles. MDS is the classic way to visualise multivariate similarity/distance data; the stress value honestly
    tells how much we can trust the map. It is used to see participant grouping and profile similarity.

====================================================================================

#72  UMAP
    file: 72_umap.xlsx
  >> SCENARIO (narration):
    Imagine we administered a 25-item behavioral inventory to 200 participants,
    and we suspect there are hidden subgroups in how people behave that aren't
    obvious from any single item. We also recorded a coarse classification,
    lower_type, with five categories S1 through S5, that we want to overlay
    afterward to see if the structure lines up. Our goal is to take that high-
    dimensional set of behavior_01 through behavior_25 and squeeze it into a
    two-dimensional embedding that preserves the local neighborhood structure.
    UMAP is the right tool because it handles many correlated items, reveals
    nonlinear clusters that linear methods miss, and gives us a map we can color
    by lower_type to check whether the emergent groups match our labels.
  >> VARIABLE SELECTION:
    - Features (high-dimensional input): behavior_01 ... behavior_25 (all 25 behavior items)
    - Color / label overlay (optional): lower_type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    UMAP dimension reduction (25 behavior items -> 2D)
    n_neighbors and min_dist balance local/global structure

>> COMMENTARY (narration):
    UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure
    better and is faster. We projected participant profiles of 25 behavior items into a two-dimensional map; similar
    behavior profiles cluster. The n_neighbors parameter tunes the local-global balance, and min_dist the tightness
    of clusters. UMAP is increasingly preferred for discovering hidden group structure in high-dimensional behavior/
    survey data. It is used for visualisation; the clustering pattern is interpreted, not the axis values.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach.xlsx
  >> SCENARIO (narration):
    Suppose we developed a 21-item self-report scale and administered it to 200
    participants, and before we use it in any analysis we need to know whether
    the items hang together as a reliable measure of a single construct. Each
    item, item_01 through item_21, is scored on a numeric rating. The question
    is simple but crucial: are these items measuring the same underlying thing
    consistently? Cronbach's Alpha is the standard choice here because it
    quantifies internal consistency across a set of items, tells us the overall
    reliability coefficient, and shows how alpha would change if any item were
    dropped.
  >> VARIABLE SELECTION:
    - Items (scale items): item_01 ... item_21 (all 21 items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.972 (excellent internal consistency)   21 items

>> COMMENTARY (narration):
    We tested the internal consistency of a 21-item psychological scale with Cronbach's alpha. Alpha = 0.97 is
    excellent: the 21 items consistently measure the same underlying concept, and participants answer coherently
    across items. Showing a scale's reliability before analysing it is essential -- low alpha signals the items do
    not measure the same thing and the total score may be meaningless. Above 0.70 is acceptable, above 0.90 excellent.
    It is the first quality step of any psychological study developing a scale. (A very high alpha can also hint at
    item redundancy.)

====================================================================================

#74  Likert Analysis
    file: 74_likert.xlsx
  >> SCENARIO (narration):
    Imagine we ran a questionnaire on 220 participants that taps three
    psychological constructs measured with Likert-type items: five motivation
    items, five self-efficacy items, and five anxiety items. We want to
    summarize how respondents distributed across the response options for each
    item and visualize the agreement patterns rather than just reporting raw
    means. Likert Analysis is the appropriate approach because it treats these
    ordered rating items properly, produces per-item frequency and percentage
    breakdowns, and lets us see the response distribution across the motivation,
    self-efficacy, and anxiety blocks at a glance.
  >> VARIABLE SELECTION:
    - Items (Likert items): motivation_1 ... motivation_5
    - Items (Likert items): self_efficacy_1 ... self_efficacy_5
    - Items (Likert items): anxiety_1 ... anxiety_5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.769 (acceptable)   15 items (3 dimensions: motivation/self-efficacy/anxiety)

>> COMMENTARY (narration):
    We examined the overall reliability and item statistics of a 15-item, three-dimensional (motivation,
    self-efficacy, anxiety) Likert scale. Cronbach's alpha = 0.77 is "acceptable" -- the scale is generally
    consistent, but because three different dimensions are combined, a single alpha can come out a bit low
    (sub-dimensions should be analysed separately). Likert analysis gives item means, distributions and item-total
    correlations to show whether the scale is sound. It is used for the basic quality control of attitude/motivation
    scales in psychology.

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa.xlsx
  >> SCENARIO (narration):
    Suppose we built an 18-question attitude survey and collected responses from
    250 participants, but we don't yet know how many underlying dimensions these
    questions actually capture. The items q01 through q18 were written to cover
    what we think are a few related themes, yet we want the data itself to
    reveal the latent structure. Exploratory Factor Analysis is the natural fit
    because we have no fixed hypothesis about which item loads on which factor;
    EFA will extract the latent factors, show us the loadings, and help us
    decide how many meaningful dimensions lie behind the 18 items.
  >> VARIABLE SELECTION:
    - Items (observed variables): q01 ... q18 (all 18 questionnaire items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO (overall) = 0.908 (Excellent)   -> factor analysis very appropriate
    18 items, expected 3 factors

>> COMMENTARY (narration):
    We explored how many hidden dimensions (factors) underlie 18 scale items with Exploratory Factor Analysis. KMO =
    0.91 is "excellent": the data are very suitable for factor analysis, with high shared variance among items. The
    items group into three factors as expected. EFA reduces many items to a few meaningful dimensions, simplifying
    the scale and revealing the underlying structures. It is a core tool of psychological scale development and
    psychometric validation; KMO and Bartlett's test are reported as the pre-check of suitability.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc.xlsx
  >> SCENARIO (narration):
    Imagine three counselors independently rated the same 50 participants on a
    numeric scale, and before we trust these ratings we need to know how
    consistent the raters are with one another. Each participant has a score
    from counselor_a, counselor_b, and counselor_c. The question is whether
    different counselors produce essentially interchangeable ratings of the same
    person. Intraclass Correlation is exactly the right measure here because it
    quantifies the agreement among multiple raters on continuous scores,
    distinguishing genuine between-subject differences from rater inconsistency.
  >> VARIABLE SELECTION:
    - Raters (measurements): counselor_a
    - Raters (measurements): counselor_b
    - Raters (measurements): counselor_c

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.961   (Excellent)   3 counselors   95% CI ~ (0.94 , 0.98)

>> COMMENTARY (narration):
    We assessed how consistent three counselors are in scoring the same participants with ICC. ICC = 0.96 is
    "excellent": inter-counselor agreement is very high -- whichever counselor scores, the result is similar. This is
    critical in psychological assessment: if scoring varied from counselor to counselor, assessments would be
    unreliable. ICC is the standard for measuring inter-rater reliability on continuous measurements (scores, ratings)
    (Cohen's Kappa is for categorical data, ICC for continuous). It is essential for clinical-assessment reliability.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa.xlsx
  >> SCENARIO (narration):
    Suppose prior theory tells us that wellbeing has three distinct facets,
    self-related, psychological, and social, and we wrote four items for each,
    giving twelve observed indicators measured on 280 participants. Unlike an
    exploratory approach, here we already have a specific measurement model in
    mind: self_1 to self_4 load on one factor, psychological_1 to
    psychological_4 on a second, and social_1 to social_4 on a third.
    Confirmatory Factor Analysis is the correct method because it tests this
    pre-specified three-factor structure directly, evaluating how well our
    hypothesized loadings fit the observed data through goodness-of-fit indices.
  >> VARIABLE SELECTION:
    - Factor 1 indicators (self): self_1, self_2, self_3, self_4
    - Factor 2 indicators (psychological): psychological_1, psychological_2, psychological_3, psychological_4
    - Factor 3 indicators (social): social_1, social_2, social_3, social_4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 1.00   RMSEA = 0.000   (excellent fit)
    3 factors: self + psychological + social

>> COMMENTARY (narration):
    While EFA explores hidden structure, CFA tests a pre-specified theory: here we tested whether three distinct
    well-being dimensions -- "self", "psychological" and "social" -- fit the data. The fit is excellent (CFI = 1.00,
    RMSEA = 0.000): the items really load onto these three dimensions, the hypothesised scale structure is fully
    confirmed. CFA is part of structural equation modelling (SEM) and is the strongest way to prove scale validity.
    It is the standard method for demonstrating the validity of multidimensional constructs like psychological
    well-being and self-concept.

====================================================================================

#78  Survey Means
    file: 78_survey_means.xlsx
  >> SCENARIO (narration):
    Imagine a national survey of 350 participants where we measured a verbal
    ability score and, importantly, recorded a sampling weight for each
    respondent so the sample reflects the true population. Participants come
    from four regions: east, Marmara, G.east, and inner anatolia. We want an
    unbiased estimate of the average verbal score, both overall and broken down
    by region, that properly accounts for unequal selection probabilities.
    Survey Means is the right procedure because it computes weighted means and
    their correct standard errors using the design weight, which a naive average
    would get wrong.
  >> VARIABLE SELECTION:
    - Variable (continuous outcome): survey_verbal
    - Weight: weight
    - Domain / by group (optional): region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Verbal score weighted mean = 458.99   SE = 3.99   95% CI (451.14 , 466.83)   CV 0.87%
    Stratified + weighted design (Taylor SE)

>> COMMENTARY (narration):
    Large-scale social surveys select participants not at random but with a stratified, weighted design. A simple
    mean would be wrong for such data; complex-sample analysis gives the correct mean and standard error (Taylor
    linearization) by accounting for the design weights and stratification. The verbal-score weighted mean is 459,
    with a very narrow confidence interval (CV 0.87%). National social surveys (values survey, household survey) are
    always designed this way; survey methods are essential for population-generalisable, unbiased estimation.

====================================================================================

#79  Survey Frequencies
    file: 79_survey_freq.xlsx
  >> SCENARIO (narration):
    Suppose we surveyed 400 respondents about their institutional intention,
    recorded as maybe, probable, none, or certain, and we also know whether each
    respondent lives in a rural or urban area. Because this was a weighted
    survey, each respondent carries a sampling weight. We want to report what
    proportion of the population falls into each intention category, with
    design-correct confidence intervals, and compare those proportions across
    the rural and urban groups. Survey Frequencies is the appropriate analysis
    because it produces weighted category proportions and their standard errors
    for a categorical variable under a complex sampling design.
  >> VARIABLE SELECTION:
    - Categorical variable: institution_intention
    - Weight: weight
    - Domain / by group (optional): region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    "certain" proportion = 0.095 (SE 0.015)   |   "maybe" proportion = 0.363 (SE 0.024)
    Institution intention -- stratified + weighted design

>> COMMENTARY (narration):
    We estimated the proportion of institution-intention categories among participants, accounting for the complex
    sample design. Those saying "maybe" are 36%, "certain" 10% (with design-based standard errors). A simple
    percentage would give the wrong standard error by ignoring stratification and weights; survey frequency analysis
    produces a confidence interval reflecting the true uncertainty. It is the correct way to estimate category
    proportions in a population (intention, attitude, behavior distribution) in a design-consistent manner -- the
    basis of social-policy monitoring.

====================================================================================

#80  Survey Totals
    file: 80_survey_total.xlsx
  >> SCENARIO (narration):
    Imagine we sampled 280 communities to estimate the total number of
    counselors across a population, where communities fall into three size
    strata labeled M, L, and S. Each sampled community has a counselor_count and
    a sampling weight reflecting how many communities it represents. Our goal is
    not an average but a population total, the estimated overall count of
    counselors, with a proper variance estimate, and we'd like it reported by
    stratum. Survey Totals is the correct tool because it scales the observed
    counts up to a weighted population total and computes its design-based
    standard error.
  >> VARIABLE SELECTION:
    - Variable (to total): counselor_count
    - Weight: weight
    - Stratum / by group (optional): stratum

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total counselor count = 67,866   SE = 1,315   95% CI (65,278 , 70,454)
    Stratified design (Taylor SE)

>> COMMENTARY (narration):
    We estimated the total number of counselors in a region using the sampling design and weights: about 67,866 (95%
    CI 65,278 - 70,454). Scaling from the sampled communities up to the whole population requires correct weighting --
    each sampled unit carries weight proportional to the population fraction it represents. Survey total analysis does
    exactly this and expresses uncertainty with the standard error. It is the official-statistics method used to
    estimate population totals (total staff, total households, total population) from a sample.

====================================================================================

#81  Survey Regression
    file: 81_survey_reg.xlsx
  >> SCENARIO (narration):
    Suppose a national well-being survey sampled 320 adults across two regions,
    and we want to know what predicts their composite survey score. We have each
    respondent's age, their reported noise/sound exposure level, and the region
    they live in (east or Marmara), and crucially every respondent carries a
    sampling weight so that the sample reflects the true population structure.
    We use survey-weighted linear regression here because an ordinary regression
    would ignore the unequal selection probabilities and give biased estimates
    and standard errors; incorporating the weight lets us draw conclusions that
    generalize to the population the survey was designed to represent.
  >> VARIABLE SELECTION:
    - Dependent variable: survey
    - Predictors: age
    - Predictors: sound
    - Predictors (categorical): region
    - Weight: weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.185   age coefficient = 14.12   p < .001 ***
    Verbal score ~ age + socio-economic   (stratified + weighted)

>> COMMENTARY (narration):
    We modelled verbal score from age and socio-economic level, accounting for the complex sample design (strata,
    weights). The age effect is significant (coefficient 14.12, p < .001). The model explains 18% of the variance --
    modest, so other factors also drive verbal score. Design-based standard errors are more accurate than the (often
    too optimistic) errors of simple OLS. When estimating relationships from national social survey data (where the
    sample is not random), survey regression is necessary for valid inference; otherwise p-values would mislead.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic.xlsx
  >> SCENARIO (narration):
    Imagine we ran a weighted social survey of 320 participants and want to
    model whether someone passed an institutional milestone, recorded as
    institution_passed (0/1). We think their well-being score, their exposure to
    noise, and their region (east versus Marmara) all matter, but because the
    survey used a complex sampling design we must account for the design weight
    when estimating the odds. We use survey-weighted logistic regression because
    the outcome is binary and the data are not a simple random sample; this
    gives us population-level odds ratios with correct, design-adjusted standard
    errors.
  >> VARIABLE SELECTION:
    - Dependent variable (binary): institution_passed
    - Predictors: wellbeing
    - Predictors: sound
    - Predictors (categorical): region
    - Weight: weight

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    wellbeing: OR = 1.053   p < .001 ***
    Institution attainment ~ well-being + socio-economic   (stratified + weighted)

>> COMMENTARY (narration):
    Predicting a participant's probability of institution attainment from well-being, we handled both the binary
    outcome (logistic) and the complex sample design together. The odds ratio for well-being is 1.053 (p < .001): a
    one-unit rise in well-being raises the odds of attainment about 5%. Because the design weights and stratification
    are accounted for, the OR and confidence interval are design-consistent and valid. Survey logistic regression is
    the correct way to model binary outcomes (attainment, success, yes/no) from complex social surveys.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam.xlsx
  >> SCENARIO (narration):
    Let's say we are studying how the amount of weekly activity relates to a
    psychological outcome score in 200 participants, but we suspect the
    relationship is not a straight line: very low and very high activity might
    both behave differently than the middle range. Instead of forcing a linear
    fit, we want the data to reveal the shape of the curve itself. We use a
    Generalized Additive Model because it fits a flexible smooth function of
    activity rather than a single slope, letting us see and test possible non-
    linear, curved effects on the score.
  >> VARIABLE SELECTION:
    - Dependent variable: score
    - Smooth predictor: activity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (explained) = 0.756
    Score ~ s(activity)  (smooth function)

>> COMMENTARY (narration):
    GAM is a flexible generalisation of linear regression: it models a variable's effect with "smooth curves"
    (splines) rather than a straight line. Activity's effect on score is often nonlinear -- it may show diminishing
    returns beyond a point. GAM learns this curved pattern from the data without assuming its shape in advance; the
    model explains 76% of the variance. It preserves interpretability while adding flexibility: we can read from the
    curve where activity helps most. For nonlinear but interpretable relationships in psychology, it is an ideal
    method.

====================================================================================

#84  Discriminant Analysis (LDA/QDA)
    file: 84_diskriminant.xlsx
  >> SCENARIO (narration):
    Suppose we have 105 students classified into three cohorts of academic
    standing - low, medium, and high - and we measured several psychological and
    performance variables: math and science scores, well-being, withdrawal
    tendency, and weekly activity. We want to know which combination of these
    measures best separates the three cohorts and whether we can correctly
    classify a new person into the right group. We use Discriminant Analysis
    (LDA/QDA) because the outcome is a categorical group with more than two
    levels and the predictors are continuous, making it ideal for finding the
    discriminant functions that maximize between-group separation.
  >> VARIABLE SELECTION:
    - Grouping variable: cohort
    - Predictors: math
    - Predictors: science
    - Predictors: wellbeing
    - Predictors: withdrawal
    - Predictors: activity

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (cohort/group classification)
    Separation from 5 variables (math, science, well-being, withdrawal, activity)

>> COMMENTARY (narration):
    Discriminant analysis finds the linear combinations of variables that best separate the groups (cohorts) -- a bit
    like a classification-focused PCA. Accuracy is 100%: five variables separate the groups perfectly. Discriminant
    analysis both classifies and answers "which variables matter most for the separation" (discriminant function
    loadings). It resembles logistic regression but is the classic choice for multiple groups when variables are
    normally distributed. In the social sciences it is used in group classification and for assigning new participants
    to existing groups.

====================================================================================

#85  Conditional Logit (Choice Model)
    file: 85_conditional_logit.xlsx
  >> SCENARIO (narration):
    Imagine we ran a choice experiment where each of our participants faced sets
    of institution alternatives and picked one. The data are in long format:
    every choice_id groups the alternatives a person saw, the institution
    attribute describes each option (public-near, public-far, private-near,
    private-far), and we recorded each option's fee and distance in kilometers,
    with chosen flagging the one actually selected. We want to know how fee and
    distance drive these decisions. We use a Conditional Logit choice model
    because it models the probability of selecting an alternative within each
    choice set as a function of the alternatives' attributes, which is exactly
    how discrete-choice data behave.
  >> VARIABLE SELECTION:
    - Choice set ID: choice_id
    - Chosen indicator: chosen
    - Alternative attributes: fee
    - Alternative attributes: distance_km
    - Alternative attributes (categorical): institution

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (McFadden) = 1.000  (perfect)
    Institution choice ~ fee + distance (km)

>> COMMENTARY (narration):
    We examined which institution participants choose with a conditional logit model -- each participant selects one
    from a choice set. McFadden pseudo R2 = 1.0 is perfect (0.2-0.4 is already strong for choice models): fee and
    distance fully explain the choice. The model shows participants prefer nearby, lower-fee institutions. This method
    underlies "discrete choice" analysis in the social sciences and economics; it is powerful for modelling
    institution/service choice, access and cost effects.

====================================================================================

#86  Kaplan-Meier Survival Analysis
    file: 86_km.xlsx
  >> SCENARIO (narration):
    Suppose we are following 200 students over time to see how long they take to
    graduate, where duration_wave records the number of waves until the event
    and graduate marks whether graduation actually occurred (1) or the person
    was still censored (0). We also know each student's department - social,
    engineering, or health - and want to compare how quickly the groups reach
    graduation. We use Kaplan-Meier survival analysis because it handles
    censored time-to-event data properly and lets us plot and compare the
    survival curves of the three departments without assuming any particular
    distribution.
  >> VARIABLE SELECTION:
    - Time: duration_wave
    - Event: graduate
    - Group (stratum): department

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 8.4 waves   (grouped by department)
    Log-rank test compares groups

>> COMMENTARY (narration):
    We analysed participants' time in/completion of a program -- by department -- with Kaplan-Meier. Median time is
    8.4 waves. Kaplan-Meier analyses "time-to-event" data -- correctly using not-yet-event (censored) participants
    too -- and tests group differences in duration with the log-rank test. In the social sciences it is the basic
    visual and test for "time-event" data such as program completion, time to dropout, graduation; it is at the
    centre of participant-continuity analysis.

====================================================================================

#87  Cox Proportional Hazards Regression
    file: 87_cox.xlsx
  >> SCENARIO (narration):
    Let's say we want to understand which factors speed up or slow down student
    attrition over time in a cohort of 250 students. For each person we have the
    follow-up time in waves and an attrition indicator showing whether they
    dropped out, plus several covariates: age, well-being, whether they held a
    scholarship, and their level of family support. We use Cox proportional
    hazards regression because it lets us estimate how each covariate multiplies
    the hazard of dropping out at any moment, while handling censoring and
    without needing to specify the baseline hazard's shape.
  >> VARIABLE SELECTION:
    - Time: duration_wave
    - Event: attrition
    - Covariates: age
    - Covariates: wellbeing
    - Covariates: scholarship
    - Covariates: family_support

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 77 (30.8%)   Concordance = 0.650
    wellbeing: HR = 0.983   95% CI (0.973 , 0.994)   p = 0.002 **

>> COMMENTARY (narration):
    We modelled a participant's attrition risk from psychosocial variables with Cox regression. Well-being is
    significant and protective (HR = 0.983 < 1, p = 0.002): as well-being rises, attrition risk falls -- high
    well-being encourages staying in the program. Cox regression is the gold-standard survival model relating
    "time-to-event" to several predictors; it gives hazard ratios (HR). Anticipating program continuity -- for early
    intervention -- is very valuable in social services. HR < 1 is protective, HR > 1 risk-increasing.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (AFT) Model
    file: 88_aft.xlsx
  >> SCENARIO (narration):
    Imagine we are studying time-to-graduation for 200 students and, rather than
    modeling the hazard, we want to model the actual survival time directly -
    for instance, by how much a scholarship multiplies or shortens the expected
    time until graduation. The data give us duration_wave as the follow-up time,
    graduate as the event indicator, and scholarship as a covariate. We use a
    Parametric Survival (AFT) model because it assumes a specific distribution
    for the survival times and expresses covariate effects as acceleration or
    deceleration of the time scale, which yields interpretable time-ratio
    estimates.
  >> VARIABLE SELECTION:
    - Time: duration_wave
    - Event: graduate
    - Covariates: scholarship

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution = Weibull   Median survival = 6.91 waves
    Survival ~ scholarship

>> COMMENTARY (narration):
    We analysed participants' time in the program with a parametric Weibull model. While Cox is semi-parametric,
    Weibull models the whole survival curve with a specific mathematical distribution -- if the data fit it, it gives
    more efficient estimates and allows extrapolation. Median time is 6.91 waves. The AFT (accelerated failure time)
    interpretation is intuitive: how many times a factor like scholarship lengthens/shortens the time. When the time
    distribution is known, parametric models give more powerful and predictable results than Cox.

====================================================================================

#89  Competing Risks
    file: 89_competing.xlsx
  >> SCENARIO (narration):
    Suppose we follow 220 students whose studies can end in several mutually
    exclusive ways: graduation, continuation, attrition, or transfer, recorded
    in the separation variable, with duration_wave giving the time and event
    flagging whether a terminal outcome happened. The problem is that
    experiencing one outcome, like transferring, removes a student from the risk
    of the others, so a standard survival analysis would misestimate the chance
    of, say, attrition. We use a Competing Risks analysis because it correctly
    models multiple distinct event types and estimates the cumulative incidence
    of each one in the presence of the others.
  >> VARIABLE SELECTION:
    - Time: duration_wave
    - Event indicator: event
    - Event type (competing outcomes): separation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 49   |   3 different separation types (graduation, attrition, transfer)
    Separate cumulative incidence (CIF) and risk for each event type

>> COMMENTARY (narration):
    A participant can leave a program for more than one reason: graduation (positive), attrition, or transfer -- and
    these are mutually exclusive (once one happens the others cannot). Classic survival handles this incorrectly;
    competing-risks analysis produces a separate cumulative incidence for each separation type. In the data 49
    participants are still in the program (censored), the rest left for different reasons. This framework correctly
    answers questions like "which factor increases the risk of which separation type". In the social sciences, when
    there are different outcome pathways (graduation/attrition/transfer), competing-risks analysis is the correct and
    informative method.

====================================================================================

#90  Time-Dependent Cox Model
    file: 90_tvcox.xlsx
  >> SCENARIO (narration):
    Let's say we are tracking students' risk of dropping out, but well-being is
    not fixed - it changes over the follow-up period, so we recorded the data in
    start-stop interval format where each row covers a time segment from start
    to end, carries the well-being value during that segment, and an event flag
    for whether dropout occurred in it. A standard Cox model assumes covariates
    stay constant, which would be wrong here. We use a Time-Dependent Cox model
    because it allows well-being to vary across each person's follow-up
    intervals, correctly linking the current covariate value to the hazard at
    each point in time.
  >> VARIABLE SELECTION:
    - Start time: start
    - Stop time: end
    - Event: event
    - Time-dependent covariate: wellbeing

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    wellbeing: HR = 1.001   95% CI (0.983 , 1.019)   p = 0.944 (non-significant)
    Time-varying well-being covariate

>> COMMENTARY (narration):
    A participant's well-being changes from wave to wave; standard Cox assumes it constant, time-dependent Cox uses
    the current value in each time interval (start-stop format). In this model the effect of current well-being on
    attrition risk is non-significant (HR = 1.001, p = 0.944) -- in this data the momentary well-being change does not
    predict attrition timing. A non-significant result is information too. Time-dependent Cox is the correct method
    when "the current, continuously monitored condition matters, not just the baseline"; it is the standard for
    modelling time-varying factors in longitudinal psychological monitoring.

====================================================================================

#91  Survey-Weighted Cox Regression
    file: 91_survey_phreg.xlsx
  >> SCENARIO (narration):
    In this study we follow 300 university students to see how long they stay
    enrolled before dropping out, but our sample comes from a complex survey
    design rather than a simple random draw. Each participant carries a sampling
    weight that reflects their probability of selection across communities, so
    ignoring it would bias our estimates. We model the time-to-attrition outcome
    (duration_wave) with attrition as the event indicator, comparing risk across
    regions A, B and C while honoring the survey weights and community
    clustering. Survey-weighted Cox regression is the right tool here because it
    estimates population-level hazard ratios that remain valid under the
    weighted, clustered sampling scheme.
  >> VARIABLE SELECTION:
    - Time (duration): duration_wave
    - Event: attrition
    - Predictor (covariate): region
    - Weight: weight
    - Cluster: community_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.50   Weighted/clustered survival (region covariate, community cluster)
    Attrition ~ region

>> COMMENTARY (narration):
    We combined survival analysis with a complex sample design: participants come from a population sampled with
    weights and clustered by communities. Survey-PHREG runs the Cox model with design weights and cluster-robust
    standard errors, so the hazard ratios and confidence intervals generalise to the population. Concordance 0.50
    means region alone is a weak discriminator in this model. When national social surveys have "time-to-event" data,
    ignoring the design biases the estimates; survey-PHREG provides valid inference.

====================================================================================

#92  Interval-Censored Survival Analysis
    file: 92_interval.xlsx
  >> SCENARIO (narration):
    Here we want to know when students actually complete a transition, but we
    never observe the exact moment. Instead, each of the 60 participants was
    checked at two assessment waves, giving us only a lower and upper bound on
    when the event happened. The columns left_censor and survival_censor bracket
    that unknown event time, and we compare two departments, A and B. Interval-
    censored survival analysis is appropriate because the event time falls
    inside a known window rather than being precisely recorded, and ordinary Cox
    or Kaplan-Meier methods would mishandle that uncertainty.
  >> VARIABLE SELECTION:
    - Left bound: left_censor
    - Right bound: survival_censor
    - Group: department

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 48   Median survival = 12.0 waves   (Turnbull algorithm)
    Event time known within [lower, upper] interval

>> COMMENTARY (narration):
    In social monitoring we rarely know the exact time of an event: we observe participants in periodic waves, find
    them in one state at one wave and another at the next -- the event happened somewhere between the two
    measurements. Interval-censored survival (Turnbull algorithm) handles exactly this uncertainty, avoiding fixing
    the event to an arbitrary point. Median survival is 12 waves. By the nature of periodic panel measurement (annual
    wave, periodic measurement), event times are always within an interval; forcing this data to a single point
    biases the result. The interval-censored method matches the reality of panel data.

====================================================================================

#93  Frailty Cox Model
    file: 93_frailty.xlsx
  >> SCENARIO (narration):
    In this dataset 120 students are nested within communities, and we suspect
    that unobserved community-level factors make some neighborhoods more prone
    to attrition than others. We track time-to-dropout (duration_wave) with
    attrition as the event, and we want to estimate the effect of clinical_group
    (K1 through K4) on the dropout hazard while accounting for that shared
    community frailty. A frailty Cox model fits because it adds a random effect
    for community_id, capturing within-community correlation that a standard Cox
    model would wrongly ignore.
  >> VARIABLE SELECTION:
    - Time (duration): duration_wave
    - Event: attrition
    - Predictor (covariate): clinical_group
    - Frailty (cluster): community_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.50   Community-clustered shared frailty (random effect)
    Attrition ~ clinical group  +  (community frailty)

>> COMMENTARY (narration):
    Participants are clustered within communities; participants in the same community share unmeasured common factors
    (social environment, resources, culture) and carry similar attrition risk. Frailty Cox adds a shared "frailty"
    (random effect) for each community to model this clustering -- the mixed-model version of survival analysis.
    Accounting for community-level hidden differences gives both correct standard errors and information on "how much
    heterogeneity exists between communities". It is the correct method for clustered social survival data
    (communities, neighbourhoods).

====================================================================================

#94  Time Series Description
    file: 94_ts.xlsx
  >> SCENARIO (narration):
    Before modeling anything fancy, we simply want to understand a 60-point
    monthly series of a behavioral record collected over time. Each row is a
    date with one observed record value, and we need a first look at its level,
    spread, range and overall movement. Time series description is the natural
    starting point because it summarizes the series, checks for gaps and gives
    the descriptive picture we need before deciding on decomposition or
    forecasting.
  >> VARIABLE SELECTION:
    - Time: date
    - Value: record

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.858 -> NOT stationary   Trend = increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly record series. The ADF test does not find it stationary (p = 0.858) -- the mean changes over
    time, there is a trend (increasing) and seasonality is present. This pre-diagnosis is critical: most time-series
    models (like ARIMA) require stationarity, so differencing/transformation may be needed first. Time-series analysis
    reveals the structure of long monitoring data -- trend, seasonality, stationarity -- and tells which modelling
    steps are required. It is the foundation for analysing how social phenomena (applications, registrations, event
    counts) evolve over time.

====================================================================================

#95  STL Decomposition
    file: 95_stl.xlsx
  >> SCENARIO (narration):
    We have a long daily series of motivation scores spanning three years, 1095
    observations in all, and we suspect it contains both a slow underlying trend
    and a repeating seasonal rhythm. To separate these pieces we want to split
    the motivation series into its trend, seasonal and remainder components. STL
    decomposition is ideal here because it robustly pulls apart trend and
    seasonality from a single long time series, letting us see the structure
    that raw daily values obscure.
  >> VARIABLE SELECTION:
    - Time: date
    - Value: motivation

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Period (seasonal length) = 7 (daily -> weekly cycle)
    Series decomposed into trend + seasonal + residual components

>> COMMENTARY (narration):
    STL (Seasonal-Trend decomposition using Loess) splits a daily motivation series into three components: long-term
    trend, a recurring cycle (7-day = weekly) and the remaining residual. The weekly period is meaningful --
    motivation fluctuates regularly across the week. This decomposition clarifies "is motivation generally rising, or
    just fluctuating weekly". STL is the most intuitive way to interpret seasonal/cyclical psychosocial series,
    separating trend from cycle so each can be assessed separately.

====================================================================================

#96  ARIMA Forecasting
    file: 96_arima.xlsx
  >> SCENARIO (narration):
    In this study we have 120 monthly counts of graduations and we want to
    forecast where the series is heading over the coming periods. The graduate
    values show autocorrelation over time that a simple average cannot capture.
    ARIMA forecasting is appropriate because it models the autoregressive and
    moving-average structure of the single graduate series and produces
    principled forecasts with confidence intervals.
  >> VARIABLE SELECTION:
    - Time: date
    - Value: graduate

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model fitted   AIC = 2512.56
    Monthly graduation/completion count -- future forecast

>> COMMENTARY (narration):
    We fitted an ARIMA model to a monthly graduation/completion-count series and forecast the future. ARIMA combines
    the series' own past values (AR), past errors (MA) and differencing (I, to make it stationary). With AIC = 2513
    the best model was selected. Forecasts rest on the observed trend and autocorrelation structure and come with an
    uncertainty band. Time-series forecasting matters for social-program planning: anticipating future completion/
    application counts enables resource and capacity planning. ARIMA is the classic forecasting method for social
    series with seasonality and trend.

====================================================================================

#97  Exponential Smoothing (Holt-Winters)
    file: 97_ets.xlsx
  >> SCENARIO (narration):
    Here we track 84 months of application volume to a counseling program and we
    want short-term forecasts that respect both the trend and the seasonal
    swings in demand. The application series rises and falls in a recurring
    yearly pattern. Exponential smoothing with Holt-Winters is the right choice
    because it explicitly models level, trend and seasonality, giving smooth,
    adaptive forecasts for this seasonal series.
  >> VARIABLE SELECTION:
    - Time: date
    - Value: application

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method = Holt-Winters (additive seasonal, period 12)   AIC = 965.83
    Monthly application count forecast

>> COMMENTARY (narration):
    Holt-Winters exponential smoothing estimates the level + trend + seasonality components in a weighted way (giving
    more weight to the recent past). It captured the seasonal pattern of monthly applications (AIC = 966).
    Applications typically show strong seasonality; Holt-Winters is often very successful for such series and is more
    intuitive than ARIMA. In social-service demand forecasting (applications, registrations, participation) it is an
    easy-to-build, interpretable alternative, especially preferred for regularly seasonally fluctuating variables.

====================================================================================

#98  Mann-Kendall Trend and Sen's Slope
    file: 98_mann_kendall.xlsx
  >> SCENARIO (narration):
    We have 40 years of an annual verbal achievement indicator and the key
    question is whether scores have been trending up or down over the decades.
    Because such social indicators are often non-normal and may move
    monotonically rather than linearly, we use the Mann-Kendall trend test
    together with Sen's slope. This pairing is suitable because Mann-Kendall
    non-parametrically detects whether a monotonic trend exists and Sen's slope
    robustly estimates its magnitude per year.
  >> VARIABLE SELECTION:
    - Time: year
    - Value: verbal_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   Trend = INCREASING
    Verbal score long-term trend

>> COMMENTARY (narration):
    We tested whether verbal scores show a significant long-term trend with the Mann-Kendall test -- a nonparametric,
    outlier- and distribution-robust trend test. The result is highly significant (p < .001) and the direction is
    increasing: verbal score is rising consistently over the years. This may reflect the positive cumulative effect of
    educational/social policies. Mann-Kendall + Sen's slope is the gold standard for detecting long-term social trends
    -- from demography to education; Sen's slope gives the trend magnitude without being affected by outliers.

====================================================================================

#99  Anomaly Detection
    file: 99_anomali.xlsx
  >> SCENARIO (narration):
    In this dataset of 200 participants we want to flag unusual cases whose
    psychological profile departs sharply from the rest of the sample. We
    examine wellbeing, activity, absence and score together, looking for records
    that are jointly atypical rather than odd on any single measure. Anomaly
    detection is appropriate because it scans the multivariate space of these
    variables to identify outliers that could signal data-entry errors or
    genuinely at-risk individuals.
  >> VARIABLE SELECTION:
    - Features: wellbeing
    - Features: activity
    - Features: absent
    - Features: score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 18   (IQR-based multivariate outlier detection)
    Extraordinary profiles in participant records

>> COMMENTARY (narration):
    We found the extraordinary profiles in participant records (well-being, activity, absence, score) with anomaly
    detection -- 18 participants were flagged as deviating from the normal pattern. These anomalies can point to real
    situations: an extraordinary psychosocial profile, a case needing special support, or a data-entry error. Anomaly
    detection automatically picks out the few special cases that need attention from large social databases. In
    psychology it is extremely practical for early warning, special referral and data-quality control.

====================================================================================

#100  Variance Components Analysis
    file: 100_varcomp_3level_h2.xlsx
  >> SCENARIO (narration):
    Here we have a three-level nested design: 100 repeated measurements taken on
    lower units, which are themselves grouped within four upper units (L_A
    through L_D). We want to know how much of the total variation in the value
    outcome sits at the upper-unit level, the lower-unit level and the
    measurement level. Variance components analysis is the right approach
    because it decomposes the outcome's variance across these nested random
    levels, which is exactly what we need to estimate reliability and intraclass
    structure.
  >> VARIABLE SELECTION:
    - Outcome (value): value
    - Level 1 (upper): upper_unit
    - Level 2 (lower): lower_unit
    - Level 3 (measurement): measurement_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper unit = 12.1%   lower unit (nested in upper) = 40.9%   residual = 47.0%
    Nested variance components (REML)

>> COMMENTARY (narration):
    We partitioned the variability of a measurement by which level -- upper unit (e.g. region), lower unit (community)
    or individual -- it comes from with variance-components analysis. The result: 40.9% of the variance is at the
    lower unit (community) level, 12.1% at the upper unit (region) level, and 47% individual/residual variation. So
    the largest structural source is the community level; between-region differences are relatively small. This is a
    critical inference in heritability (h2) and multi-level social-design studies: it shows where the variability
    concentrates and at which level (community, region or individual) an intervention should be targeted.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old.xlsx
  >> SCENARIO (narration):
    Imagine a counseling center has just redesigned its support program and
    wants to know whether the new format genuinely improves participants'
    outcomes compared to the old one. In this dataset 80 participants are split
    into two groups, 'old' and 'new', and each provides a continuous outcome
    score on the value variable. Instead of a classic significance test, we use
    a Bayesian independent-samples t-test so we can quantify the evidence
    directly: how much more likely the data are under a real difference between
    the new and old groups versus no difference. This is ideal here because the
    center wants to talk about the strength of evidence and a credible interval
    for the effect, not just a yes-or-no p-value.
  >> VARIABLE SELECTION:
    - Grouping factor (2 levels): group
    - Test variable: value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 3.36e+46 (decisive evidence)   Cohen's d = 3.94
    New method vs old method -- score

>> COMMENTARY (narration):
    We tested whether a new intervention method differs from the old with a Bayesian t-test. The Bayes Factor BF10 =
    3.36e+46 -- astronomically large, "decisive evidence": the data support the hypothesis of a difference quadrillions
    of times more than no difference. Cohen's d = 3.94 makes the effect enormous. Unlike a classic p-value, the Bayes
    Factor directly measures the strength of evidence for both H1 and H0 and distinguishes "absence of evidence" from
    "evidence of absence". The new method's effect is not just statistical but practically overwhelming -- the
    Bayesian framework shows this convincingly.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_BF10.xlsx
  >> SCENARIO (narration):
    Suppose a sociologist suspects that two psychological measures tend to move
    together, but wants to avoid over-interpreting a borderline correlation.
    Here we have 80 participants, each with two continuous scores, X_variable
    and Y_variable, and the question is simply whether they are genuinely
    related. A Bayesian correlation is the right tool because it returns a Bayes
    factor (BF10) that weighs the evidence for a real association against the
    null of no association, along with a posterior credible interval for the
    correlation. This lets us say not only how strong the relationship is, but
    how convincing the evidence actually is.
  >> VARIABLE SELECTION:
    - First variable: X_variable
    - Second variable: Y_variable

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.780 (very strong)   BF10 = 4.50e+14 (decisive evidence)
    Relationship between two variables

>> COMMENTARY (narration):
    We assessed the relationship between two variables with Bayesian correlation: r = 0.78 is very strong and the
    Bayes Factor BF10 = 4.50e+14 makes the evidence for the relationship's existence decisive. Unlike classic
    correlation, Bayesian correlation expresses the strength of the relationship with a probability distribution and
    an evidence factor -- also answering "how sure are we". Such a large BF10 says the probability that the
    relationship is coincidental is vanishingly small. It is a robust framework for quantifying strong links between
    psychosocial variables and reporting the strength of evidence.

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_2way.xlsx
  >> SCENARIO (narration):
    Picture a social-psychology experiment with a two-factor design: one factor
    with three conditions, A, B and C, and a second factor with two levels, X
    and Y. Each of the 120 participants gives a continuous outcome on the value
    variable, and we want to understand both main effects and whether the two
    factors interact. We run a Bayesian two-way ANOVA so we can compare
    competing models, for example main-effects-only versus a model with the
    interaction, and read off Bayes factors that tell us which structure the
    data support. This is more informative than a traditional ANOVA when we care
    about evidence for or against each effect, including a possible interaction.
  >> VARIABLE SELECTION:
    - First factor: factor1
    - Second factor: factor2
    - Dependent variable: value

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor1: BF10 = 76,099 (decisive evidence)
    Two-factor design -> score

>> COMMENTARY (narration):
    We examined the effect of two factors on score with Bayesian ANOVA. The first factor's effect is strong: Bayes
    Factor 76,099 -- "decisive evidence". Bayesian ANOVA gives a separate Bayes Factor for each effect and interaction,
    answering "which factor really matters" more informatively than classic ANOVA -- and can even provide "evidence
    of absence" for non-significant effects. It is a powerful way to evaluate factor effects with evidence in social/
    psychological experiments; its ability to confirm a "no effect" result is especially valuable.

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_LMM.xlsx
  >> SCENARIO (narration):
    Consider a multi-site study where participants are nested within five
    groups, G1 through G5, perhaps clinics or schools, and we believe each group
    may respond a little differently to a continuous predictor. With 200
    observations, an X_covariate predictor and a Y_response outcome, we fit a
    Bayesian hierarchical (multilevel) model that lets each group have its own
    intercept and slope while borrowing strength across groups. This partial
    pooling is exactly what we want here: it gives stable group-level estimates
    even for smaller groups and produces full posterior distributions for both
    the overall effect and the between-group variation.
  >> VARIABLE SELECTION:
    - Grouping variable (random effect): group_id
    - Predictor (fixed effect): X_covariate
    - Response variable: Y_response

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group random variance = 2.90 (8.6%)   residual = 30.79 (91.4%)   ICC = 0.086
    Y ~ X covariate  +  (1 | group)   (Bayesian)

>> COMMENTARY (narration):
    We analysed nested data (observations within groups) with a Bayesian hierarchical model -- the Bayesian
    counterpart of LMM. The group-level variance is only 8.6% of the total (ICC = 0.086); most of the variability
    (91%) is within-group/individual. So in this data between-group differences are small, observations behave largely
    independently. The Bayesian hierarchical model's strength: it expresses both within- and between-group uncertainty
    with full probability distributions and balances groups with few observations via "partial pooling". It is a
    modern method for multi-centre or nested social data.

====================================================================================

#105  Spatial Lag (SAR) Model
    file: 105_spatial_sar_spatial.xlsx
  >> SCENARIO (narration):
    Imagine mapping a social outcome across 100 neighborhoods and noticing that
    high-scoring areas tend to sit next to other high-scoring areas. We have
    each unit's coordinates (lat, lon), two predictors X1 and X2, and the
    outcome Y_value. A standard regression would ignore this clustering, so we
    fit a Spatial Lag (SAR) model, which adds a spatially lagged dependent term
    to capture how a neighborhood's outcome is influenced by its neighbors'
    outcomes. This is the appropriate choice when we think the outcome itself
    spills over across space, not just the residuals.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictors: X1
    - Predictors: X2
    - Coordinates: lat
    - Coordinates: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho (spatial lag) = 0.34   z = 4.94   p < .001   Pseudo R2 = 0.87
    Y ~ X1 + X2  +  neighbour Y

>> COMMENTARY (narration):
    Modelling a spatial outcome (e.g. a regional social indicator), we accounted for spatial dependence -- the
    tendency of nearby units to be similar. The SAR (spatial autoregressive) model links a unit's outcome to its
    neighbours' outcome too (the rho parameter). Here rho = 0.34, significantly positive (z = 4.94, p < .001): the
    neighbour effect is real and moderate -- a region's state is influenced by neighbouring regions. The model
    explains 87% of the variance. Using a spatial model is critical, because ordinary regression gives spurious
    significant results when spatial autocorrelation is present. SAR is the right method for spatial social data with
    diffusion and neighbour effects.

====================================================================================

#106  Spatial Error Model
    file: 106_spatial_error_residual.xlsx
  >> SCENARIO (narration):
    Suppose we study the same 100 spatial units, again with coordinates lat and
    lon, predictors X1 and X2, and outcome Y_value, but now we suspect that
    unmeasured factors, like shared local conditions, leak across neighboring
    areas and show up in the errors rather than the outcome itself. In that case
    a Spatial Error model is appropriate: it keeps the usual regression of
    Y_value on X1 and X2 but models spatial autocorrelation in the residuals.
    This gives us correct standard errors and unbiased inference when the
    dependence is in what we have not measured.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictors: X1
    - Predictors: X2
    - Coordinates: lat
    - Coordinates: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda (spatial error) = 0.59   z = 6.38   p < .001   Pseudo R2 = 0.84
    Y ~ X1 + X2  +  spatial error structure

>> COMMENTARY (narration):
    The spatial error model addresses a different kind of spatial dependence than SAR: not a direct effect of
    neighbour values, but correlation induced through the residuals by unobserved (omitted) spatial factors. Lambda =
    0.59 (z = 6.38, p < .001) shows this spatial error structure is strong -- there are unmeasured common spatial
    factors (regional policy, socio-economic background). Which spatial model is appropriate (SAR or SEM) depends on
    whether the effect comes from neighbour values or from unmeasured factors. SEM cleans out these unmeasured spatial
    confounders to estimate the true effect of the variables more accurately.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local.xlsx
  >> SCENARIO (narration):
    Now imagine we suspect that the relationship between our predictors and the
    outcome is not the same everywhere, that X1 and X2 matter more in some
    places than others. With 110 spatially referenced units, their lat and lon
    coordinates, predictors X1 and X2, and outcome Y_value, we use
    Geographically Weighted Regression. GWR fits a separate, locally weighted
    regression around each location, producing a surface of coefficients we can
    map. This is the right approach when the goal is to reveal spatial
    nonstationarity, that is, how the effects of our predictors vary across the
    study area.
  >> VARIABLE SELECTION:
    - Dependent variable: Y_value
    - Predictors: X1
    - Predictors: X2
    - Coordinates: lat
    - Coordinates: lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.887   (local/location-specific regression)
    Spatial variation of the relationship

>> COMMENTARY (narration):
    Ordinary regression assumes a single relationship for the whole region; but this relationship can vary in space --
    a variable's effect may be strong in one region and weak in another. GWR (Geographically Weighted Regression)
    captures this spatial heterogeneity by fitting a separate local regression at each location (R2 = 0.89). The
    output is a map of coefficients: a surface showing where a variable's effect is strong and where it is weak. This
    tests "is the relationship the same everywhere" and sets local policy priorities. In social geography, GWR reveals
    what the global model hides for mapping regional inequalities and intervention priorities.

====================================================================================

#108  Nested Mixed Model (Nested LMM)
    file: 108_nested_lmm_R_P_F.xlsx
  >> SCENARIO (narration):
    Consider an experiment with a strictly hierarchical sampling design:
    participants are organized into blocks (R1-R4), within each block into upper
    groups (P1-P4), and within those into lower subgroups (F01-F06), with a
    continuous outcome value. Because the lower group is meaningful only inside
    its upper group, and the upper group only inside its block, these grouping
    factors are nested rather than crossed. We fit a Nested Mixed Model with
    nested random effects so that variance is correctly partitioned across the
    block, upper-group and lower-group levels, giving valid estimates of the
    fixed effect and of how much variation sits at each level of the hierarchy.
  >> VARIABLE SELECTION:
    - Response variable: value
    - Random effect (outer): block
    - Random effect (nested in block): upper_group
    - Random effect (nested in upper_group): lower_group_no

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    upper group (P): F = 27.44, p < .001 ***   |   lower x upper (RxP): F = 14.34, p < .001 ***
    block [within upper group] (F): F = 2.20, p = 0.023 *

>> COMMENTARY (narration):
    In this design the measurements are nested: upper group > block > lower unit. The nested mixed model tests each
    level's contribution to the variance separately. The results are significant at all levels: between-upper-group
    differences are largest (F = 27.44), but there are real differences at the block and lower-unit levels too.
    Mixing up the levels in nested data (e.g. ignoring the upper group) produces spurious significance or wrong
    standard errors. Nested LMM is the correct framework for hierarchical social designs (region/community/individual,
    institution/group/person), separating the genuine contribution of each scale.

====================================================================================

#109  Crossed Mixed Model (Crossed LMM)
    file: 109_crossed_lmm_A_B.xlsx
  >> SCENARIO (narration):
    Picture a design where every level of one factor is observed together with
    every level of another, for instance each of four conditions, A1-A4,
    measured under each of four contexts, B1-B4, giving a continuous outcome
    value across 128 observations. Here factor_a and factor_b are fully crossed,
    not nested, so we use a Crossed Mixed Model that treats both as random
    effects simultaneously. This lets us separate the variability due to
    factor_a from that due to factor_b while estimating the overall mean, which
    is exactly what a crossed random-effects structure is designed to handle.
  >> VARIABLE SELECTION:
    - Response variable: value
    - Crossed random effect 1: factor_a
    - Crossed random effect 2: factor_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    factor A: F = 0.38, p = 0.77 (non-significant)   |   factor B: F = 0.90, p = 0.48 (non-significant)
    A x B interaction: F = 14.67, p < .001 *** (SIGNIFICANT)

>> COMMENTARY (narration):
    Unlike nesting, here two factors are crossed: each A level is combined with each B level (fully factorial). A
    striking result: both main effects are non-significant (A: p = 0.77, B: p = 0.48), but the interaction is highly
    significant (F = 14.67, p < .001). This is a classic case of a "pure interaction": A and B alone say nothing, but
    TOGETHER -- specific A-B combinations -- they create a strong effect. Looking only at the main effects and saying
    "nothing is significant" would be a major error; the interaction changes everything. The crossed mixed model is
    the right way to capture interactions that emerge when factors are tested jointly -- a critical inference in
    social experiments.

====================================================================================

#110  KDE Density Map
    file: 110_KDE_population_density.xlsx
  >> SCENARIO (narration):
    Finally, imagine we want to see where income is concentrated across a city
    rather than just listing averages. We have 95 geocoded units with lat and
    lon coordinates and each unit's monthly_income_TL. A Kernel Density
    Estimation map smooths these point values into a continuous surface,
    revealing hotspots and gaps in the spatial distribution of income. This is
    the right visualization when the goal is to communicate where a
    socioeconomic variable is densest across space, turning scattered points
    into an intuitive density heatmap.
  >> VARIABLE SELECTION:
    - Coordinates: lat
    - Coordinates: lon
    - Weight (intensity value): monthly_income_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    95 unit points   mean monthly income = 26,569 TL
    Population/income-density map via kernel density estimation

>> COMMENTARY (narration):
    On the Map tab, we render the geographic density of population/household units into a heat map with KDE (Kernel
    Density Estimation). KDE turns points into a smooth density surface: we can see where settlement is dense and
    where there are gaps. The distribution of 95 units shows how the population/income concentrates geographically.
    This is used both to detect dense regions and service gaps. It is the basic mapping tool for visualising spatial
    social data; in social policy and urban planning it assesses the equity of population/income distribution.

====================================================================================

#111  Hexbin Map
    file: 111_Hexbin_household_distribution.xlsx
  >> SCENARIO (narration):
    Imagine we have surveyed 150 households across a city and recorded the exact
    location of each one along with how crowded its surrounding area feels,
    classified as low, medium, or high density. As social scientists we want to
    see how households are spatially distributed, where they cluster together
    and where the map thins out, without being misled by individual dots piling
    on top of each other. A hexbin map is ideal here because it divides the
    study area into equal hexagonal cells and counts how many households fall
    into each cell, turning a messy scatter of points into a clean picture of
    population concentration. We feed it the latitude and longitude of every
    household so the map can reveal the dense urban cores versus the sparse
    outskirts.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    150 household/unit points   aggregated into hexagonal cells
    Each cell shows the point density within it by colour

>> COMMENTARY (narration):
    The hexbin map aggregates many points into hexagon cells and shows each cell's density by colour -- a clear
    solution, alternative to KDE, when thousands of points overlap into an unreadable mess. We rendered 150 household
    points into a hexagonal grid to visualise where density is highest. Hexagonal cells carry less edge-bias than
    squares and represent neighbour relations more evenly. It is the practical way to turn dense spatial social data
    (household distribution, population location) into a readable density map.

====================================================================================

#112  Moran's I (Spatial Autocorrelation)
    file: 112_Morans_I_socioeconomic_autocorrelation.xlsx
  >> SCENARIO (narration):
    Suppose we have monthly income recorded for 80 neighborhood units, each
    pinned to its own geographic coordinates, and we suspect that wealth is not
    scattered randomly across the city but tends to concentrate in certain
    areas. The research question is whether rich neighborhoods sit next to other
    rich neighborhoods and poor ones next to poor ones, a pattern social
    geographers call spatial autocorrelation. Moran's I is the right tool
    because it gives us a single global statistic, ranging from negative to
    positive, that tells us whether similar income values cluster together more
    than we would expect by chance. We use latitude and longitude to define
    which units are neighbors and monthly_income_TL as the value whose spatial
    pattern we are testing.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon
    - Value variable: monthly_income_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.467   z = 8.31   p < .001
    Spatial autocorrelation of monthly income (strong positive)

>> COMMENTARY (narration):
    Moran's I is the global spatial-autocorrelation statistic measuring "do nearby regions have similar income". The
    result is strongly positive (I = 0.47, z = 8.31, p < .001): income is spatially clustered -- high-income regions
    lie near each other, and so do low ones. This is not chance; social stratification, the housing market and
    spatial segregation make neighbouring regions similar. This finding is a concrete indicator of spatial inequality
    and income segregation in sociology and is methodologically important: if spatial autocorrelation exists, ordinary
    statistics are wrong (hence the spatial models in #105-107). Moran's I is the starting diagnostic of spatial
    analysis.

====================================================================================

#113  Getis-Ord Hotspot
    file: 113_Getis_Ord_high_income_hotspot.xlsx
  >> SCENARIO (narration):
    Picture a study of 68 small neighborhood units where we know each unit's
    location and its average monthly income, and we want to pinpoint exactly
    where the affluent pockets of the city are. Unlike Moran's I, which tells us
    only that clustering exists overall, here we want to map the precise
    locations of statistically significant hot spots of high income and cold
    spots of low income. The Getis-Ord Gi* statistic is perfect for this because
    it scans each unit together with its neighbors and flags those surrounded by
    unusually high or low values. We supply the latitude and longitude to build
    the spatial neighborhood and monthly_income_TL as the variable whose local
    concentrations we want to detect.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon
    - Value variable: monthly_income_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    27 hot spots (high-income cluster)   27 cold spots (low-income cluster)   n = 68
    Getis-Ord Gi* local statistic (z > 1.96 significant)

>> COMMENTARY (narration):
    While Moran's I reports the overall clustering of the whole region, Getis-Ord Gi* shows WHERE the hot/cold spots
    are, location by location. 27 regions are statistically significant "hot spots" (high income surrounded by high
    income -- affluent clusters), and 27 regions are "cold spots" (low-income/poverty clusters). This gives direct
    guidance for social policy: it is most efficient to focus resources and intervention on the cold-spot
    (disadvantaged) clusters. Gi* hot-spot analysis is the standard spatial method for finding "where poverty/wealth
    concentrates geographically" -- it maps spatial inequality.

====================================================================================

#114  DBSCAN Spatial Clustering
    file: 114_DBSCAN_neighborhood_clusters.xlsx
  >> SCENARIO (narration):
    Imagine we have the map coordinates of 115 neighborhoods and we want to
    discover natural communities, that is, groups of neighborhoods that sit
    physically close together, while also identifying isolated ones that do not
    belong to any cluster. We do not know in advance how many clusters there
    are, and we do not want to force every neighborhood into a group. DBSCAN is
    well suited to this problem because it forms clusters based purely on
    spatial density, requires no preset number of clusters, and automatically
    labels low-density points as noise or outliers. We give it only the latitude
    and longitude of each neighborhood so it can group them by geographic
    proximity.
  >> VARIABLE SELECTION:
    - Coordinates (latitude): lat
    - Coordinates (longitude): lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   noise (isolated) = 11   n = 115
    Neighborhood/unit locations clustered by density

>> COMMENTARY (narration):
    On the map we clustered neighborhood/unit locations by density with DBSCAN: four dense neighborhood groups were
    found, and 11 units were flagged as "isolated/noise" points belonging to no cluster. DBSCAN's strength is that it
    does not require the number of clusters in advance and can capture irregularly shaped clusters -- and
    automatically separates sparse/distant units. The four clusters show neighborhoods group in particular regions;
    the isolated units may be rural/remote settlements. This spatial clustering is a practical tool for defining
    social-service regions, planning resource distribution and assessing geographically similar neighborhoods
    together. With this we complete the full 114-analysis tour of the Psychology/Sociology package.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_anxiety_score.xlsx
  >> SCENARIO (narration):
    We follow 40 participants measured at three time levels (pre, post, followup); each belongs to one of two therapy groups (CBT / waitlist). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on anxiety score.
  >> VARIABLE SELECTION:
    - Dependent variable: anxiety_score
    - Subject ID: participant_id
    - Between-subjects factor: therapy
    - Within-subjects factor: time

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (therapy): F(1,38) = 9.40  p = 0.004  np2 = 0.198
    Within (time):  F(2,76) = 168.69  p < .001  np2 = 0.816
    Interaction:         F(2,76) = 12.43  p < .001  np2 = 0.246
    Mauchly W = 0.805  p = 0.016   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a therapy x time mixed design we analyzed anxiety score for 40 participants (120 observations). The interaction is significant (F(2,76) = 12.43, p < .001, np2 = 0.246) *** -- the two groups' change across time differs in magnitude. The between-subjects main effect (CBT vs waitlist) is F = 9.40, p = 0.004; the within-subjects main effect (pre/post/followup) is F = 168.69, p < .001. Mauchly's test p = 0.016, so sphericity is violated, so the Greenhouse-Geisser corrected within p is read. Read the interaction first: when it is significant the group effect must be interpreted separately at each time level. In psychology, the mixed design is the standard analysis for pre-post-followup therapy trials versus a waitlist control.

====================================================================================
