====================================================================================
MerQur - Peyzaj_Mimarligi - SCENARIO + RESULT + COMMENTARY (EN, MERGED)
Data: english/datasets/Peyzaj_Mimarligi/
Each analysis: SCENARIO + VARIABLE SELECTION, then RESULT (screen) + COMMENTARY.
====================================================================================

#1  Descriptive Statistics
    file: 01_descriptive_park_inventory.xlsx
  >> SCENARIO (narration):
    We inventoried 280 parks in a city. For each park we measured area, green-area
    ratio, tree count, weekly visitors and a satisfaction score. Before testing any
    hypothesis we want the overall picture of the park stock: average size, green
    fabric, use and satisfaction ranges. So we start with descriptive statistics.
  >> VARIABLE SELECTION:
    - Variables: area_m2
    - Variables: green_area_ratio
    - Variables: tree_count
    - Variables: visitor_weekly
    - Variables: satisfaction_score
    - Grouping (categorical): park_type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    n = 280 parks
    area : mean = 8,888 m2   |   green-area ratio : mean = 0.696   |   tree count : mean = 68
    weekly visitors : mean = 515   |   satisfaction : mean = 3.58

>> COMMENTARY (narration):
    We first drew the general picture of the park inventory: 280 parks, mean area ~8,900 m2, 70% green-area ratio, 68
    trees per park, 515 weekly visitors, 3.58 satisfaction out of 5. This descriptive table sets the stage for all
    the analyses that follow -- design/region comparisons, green-area-satisfaction relationships, spatial park value.
    In landscape architecture, before any inferential test, summarising the park stock's basic features (area, green
    fabric, use, satisfaction) is essential both to check data quality and to define planning priorities.

====================================================================================

#2  Normality Tests
    file: 02_normality_visitor_satisfaction_noise.xlsx
  >> SCENARIO (narration):
    We want to check whether weekly visitors, satisfaction score and noise (dB) are
    normally distributed. The validity of our later t-tests, ANOVA and correlation
    depends on this assumption, so we test each variable's distribution separately.
  >> VARIABLE SELECTION:
    - Variables: visitor_weekly
    - Variables: satisfaction_score
    - Variables: noise_dB

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    3 continuous variables tested:
    visitor_weekly (weekly visitors) : Shapiro-Wilk = 0.995  p = 0.781   KS p = 0.370   Normal
    satisfaction_score (satisfaction) : Shapiro-Wilk = 0.934  p < .001    KS p = 0.001   Not Normal
    noise_dB (noise)                  : Shapiro-Wilk = 0.803  p < .001    KS p = 0.001   Not Normal

>> COMMENTARY (narration):
    We tested whether three continuous variables -- weekly visitors, satisfaction score and noise level -- are
    normally distributed. The result splits instructively: visitor count is normal (p = 0.78), but satisfaction and
    noise deviate significantly from normality (p < .001). This is typical in landscape data -- satisfaction is an
    ordinal/bounded scale (ceiling/floor effects), and noise is right-skewed (most parks quiet, a few very noisy). The
    practical takeaway: we can safely use parametric tests (t-test, ANOVA, Pearson) on visitors; for satisfaction and
    noise, nonparametric methods (Mann-Whitney, Kruskal-Wallis) or a transform are more appropriate. The normality
    check decides, for each variable separately, which test family is appropriate.

====================================================================================

#3  One-Sample t-Test
    file: 03_one_sample_t_green_area_ratio.xlsx
  >> SCENARIO (narration):
    We investigate whether parks' mean green-area ratio differs from a 30% planning
    reference threshold. With one group and a fixed reference, the one-sample t-test
    is appropriate.
  >> VARIABLE SELECTION:
    - Test variable: green_area_ratio
    - Test value (mu): 0.30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(99) = 28.18   p < .001 ***   Cohen d = 2.818 (Large)
    Mean green-area ratio = 0.557   (test mu = 0.30, reference threshold)   H0 REJECTED

>> COMMENTARY (narration):
    We compared parks' mean green-area ratio against a 30% planning reference threshold. The result is very strong:
    mean 55.7%, far above the threshold -- t(99) = 28.18, p < .001, with an enormous effect size d = 2.82. So these
    parks exceed the reference target in green-area ratio by more than chance can explain -- a green-rich park stock.
    The one-sample t-test is the right way to compare a landscape indicator against a known standard/norm (green-area
    standard, per-capita green target); it is widely used in urban green-space planning to assess compliance with
    targets.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left): one-sided is more powerful when the direction is known beforehand.
    - Hedges g: small-sample bias-corrected Cohen's d.
    - Effect-size CI: confidence interval around d.
    - Shapiro-Wilk / K-S: normality assumption checks.
    - Descriptives: mean/SD/SE/median/min/max/skewness/kurtosis.
    - Bootstrap CI: distribution-free CI for the mean by resampling.

====================================================================================

#4  Independent Samples t-Test
    file: 04_independent_t_coast_inner_satisfaction.xlsx
  >> SCENARIO (narration):
    We ask whether coastal and inner parks differ in user satisfaction. Two
    independent groups (location) and a continuous outcome call for the independent
    samples t-test.
  >> VARIABLE SELECTION:
    - Grouping: location
    - Measure: satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t(113) = 6.654   p < .001 ***   Cohen d = 1.242 (Large)
    Satisfaction differs significantly by location (coast/inner).   H0 REJECTED

>> COMMENTARY (narration):
    We examined whether park satisfaction differs by location -- coastal vs inner -- with an independent t-test. The
    difference is highly significant and large (t(113) = 6.65, p < .001, d = 1.24): coastal parks receive markedly
    higher satisfaction than inner ones. This shows the strong contribution of view/waterfront access to user
    experience in landscape design. The independent t-test is the basic method for comparing two independent groups'
    (location, park type) satisfaction/experience scores; the large effect says location is a practically very
    important design variable.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Variance assumption: Student (equal var) / Welch (unequal var — safer) / Auto (Levene decides).
    - Effect sizes: Hedges g, Glass's delta, CLES = P(X>Y).
    - Effect-size CI; per-group Shapiro; Levene & Bartlett homogeneity.
    - Per-group descriptives; Bootstrap CI for the mean difference.

====================================================================================

#5  Paired Samples t-Test
    file: 05_paired_t_renewal_before_post.xlsx
  >> SCENARIO (narration):
    To measure a renewal project's effect we compare the same parks' weekly visitors
    before and after. The same units are measured twice (paired), so the paired t-test
    is correct.
  >> VARIABLE SELECTION:
    - Before: visitor_before
    - After: visitor_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Paired t(49) = -16.11   p < .001 ***   Cohen d_z = -2.278 (Large)
    Mean difference (renewal before − after visitors) = -74.3   H0 REJECTED

>> COMMENTARY (narration):
    We measured a park-renewal (renovation) project's effect by comparing the same parks' visitor counts before and
    after renewal with a paired t-test. The result is striking: post-renewal visitor count rose by an average of 74
    people, t(49) = -16.11, p < .001, with an enormous effect size (d_z = -2.28). The renewal made parks markedly more
    attractive. Because we measured the same parks twice, the paired test is correct -- it removes between-park
    variability and focuses only on the within-park change. It is the gold-standard way to evaluate "before/after"
    interventions (renewal, redesign) in landscape practice.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction (two-sided / right / left).
    - Effect sizes: Hedges g; d_av (standardized by the average SD).
    - Effect-size CI; pairwise correlation between the two measures.
    - Shapiro / K-S on the differences; descriptives; Bootstrap CI of the mean difference.

====================================================================================

#6  One-Way ANOVA
    file: 06_one_way_anova_design_satisfaction.xlsx
  >> SCENARIO (narration):
    We compare the effect of four design styles on user satisfaction. More than two
    groups and a continuous outcome call for one-way ANOVA.
  >> VARIABLE SELECTION:
    - Grouping: design_style
    - Measure: satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 136) = 5.7482   p = 0.001 ***   η² = 0.1125
    DECISION: H0 REJECTED (4 design styles)

>> COMMENTARY (narration):
    We compared the effect of four design styles (formal, naturalistic, modern, mixed, etc.) on satisfaction with
    one-way ANOVA. The result is significant (F(3,136) = 5.75, p = 0.001), with effect size η² = 0.11 -- 11% of
    satisfaction is explained by design style (a moderate effect). ANOVA says at least one style differs from the
    others; a post-hoc test is needed to see which. "Which design approach raises user satisfaction" is a fundamental
    question in landscape design; ANOVA is the standard way to answer it across multiple groups.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - ANOVA variant: Classic (Fisher) or Welch (robust to unequal variances).
    - Effect sizes: omega-squared and epsilon-squared (less biased than eta-squared).
    - Assumptions: Levene, Bartlett, per-group Shapiro.
    - Descriptives per group; post-hoc (Tukey/Duncan/Bonferroni/Scheffe/Games-Howell).

====================================================================================

#7  Two-Way ANOVA
    file: 07_two_way_anova_style_region.xlsx
  >> SCENARIO (narration):
    We examine satisfaction with both design style and region, and want their
    interaction (is the same style equally good in every region?). Two categorical
    factors and a continuous outcome call for two-way ANOVA.
  >> VARIABLE SELECTION:
    - Factor 1: design_style
    - Factor 2: region
    - Measure: satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    design_style (main) : F(2) = 40.89   p < .001 ***   η²p = 0.32
    region (main)       : F(2) =  7.62   p < .001 ***   η²p = 0.08

>> COMMENTARY (narration):
    We examined satisfaction with two factors -- design style and region -- together. The design-style effect is very
    strong (F = 40.89, η²p = 0.32), the region effect also significant but smaller (F = 7.62, p < .001, η²p = 0.08).
    So the main determinant of satisfaction is design; region is secondary. Two-way ANOVA's strength is seeing both
    factors' effects -- separately and jointly (interaction) -- at once. This answers "does the same design style give
    the same satisfaction in every region" -- the core question of context-sensitive landscape design adapted from
    city to city.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III): Type III for unbalanced designs with interaction (SPSS default).
    - Post-hoc (Tukey/Bonferroni/Games-Howell) for 3+ level factors.
    - Effect sizes: partial eta-squared, eta-squared, omega-squared.
    - Levene & residual Shapiro; cell and marginal means tables.

====================================================================================

#8  Repeated-Measures ANOVA
    file: 08_repeated_anova_season_visitor.xlsx
  >> SCENARIO (narration):
    We compare the same parks' visitor counts across four seasons. The same units are
    measured repeatedly (dependent observations), so repeated-measures ANOVA is correct.
  >> VARIABLE SELECTION:
    - Repeated measures: season_spring_visitor, season_summer_visitor, season_autumn_visitor, season_winter_visitor

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    F(3, 237) = 1090.49   p < .001 ***   η²p = 0.9324   (n = 80)
    DECISION: H0 REJECTED

>> COMMENTARY (narration):
    We compared the same 80 parks' visitor counts across four seasons with repeated-measures ANOVA. Because the same
    parks are measured repeatedly, observations are dependent; RM-ANOVA accounts for this within-unit correlation. The
    result is overwhelming (F(3,237) = 1090.49, p < .001, η²p = 0.93): visitor count changes very markedly across
    seasons -- strong seasonality (summer peak, winter low). Effect size 93%! RM-ANOVA is the correct design for
    seasonal-use studies where the same parks are tracked over time, and is far more powerful than treating each
    season as independent. Profiling parks' seasonal-use dynamics is valuable for maintenance and event planning.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sphericity correction: Greenhouse-Geisser when Mauchly's test is violated.
    - Mauchly's sphericity test (W, p).
    - Generalized eta-squared (ges) effect size.
    - Post-hoc pairwise (Bonferroni/Holm); descriptives per level.

====================================================================================

#9  MANOVA
    file: 09_manova_style_3DV.xlsx
  >> SCENARIO (narration):
    We test design style's effect on three correlated user-experience variables
    (satisfaction, stay duration, repeat-visit ratio) at once. Multiple related DVs
    call for MANOVA.
  >> VARIABLE SELECTION:
    - Factor: design_style
    - Dependent variables: satisfaction, stay_min, repeat_visit_ratio

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilks' Lambda = 0.386 -> F(6, 230) = 23.33   p < .001 ***   (n = 120)
    Dependent: satisfaction, stay duration, repeat-visit ratio   |   Factor: design style

>> COMMENTARY (narration):
    Here we have three correlated user-experience variables -- satisfaction, stay duration, repeat-visit ratio -- and
    one design-style factor. Running three separate ANOVAs would inflate the error rate and ignore the correlations
    among variables. MANOVA tests all three jointly: Wilks' Lambda F = 23.33, p < .001, so design style strongly
    shifts the multivariate user-experience profile. MANOVA is the right approach when a design intervention affects
    several related outcomes at once. After a significant MANOVA, follow-up univariate tests show which experience
    dimension drives the effect.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Reference test for the overall decision: Wilks / Pillai (most robust) / Hotelling-Lawley / Roy.
    - Box's M: equality of covariance matrices across groups.
    - Univariate follow-up ANOVAs (one per dependent variable).
    - Multivariate partial eta-squared; per-group descriptive means.

====================================================================================

#10  ANCOVA
    file: 10_ancova_green_intervention.xlsx
  >> SCENARIO (narration):
    We examine a greening intervention's effect on final green ratio while controlling
    for parks' baseline green ratio. ANCOVA holds the baseline (covariate) constant.
  >> VARIABLE SELECTION:
    - Dependent: green_after_pct
    - Factor: intervention
    - Covariate: green_baseline_pct

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    intervention (factor): η²p = 0.588 (very large)   p < .001 ***
    Covariate: green_baseline_pct (baseline green ratio)   |   DV: green_after_pct   H0 REJECTED

>> COMMENTARY (narration):
    When evaluating a greening intervention's effect, it is essential to control for parks' baseline green ratios --
    because if parks differ at baseline, the final green-ratio difference is misleading. ANCOVA takes the baseline
    ratio as a covariate and holds it statistically constant. The result: even after adjusting for baseline, the
    intervention groups differ very significantly in final green ratio (η²p = 0.59, very large effect). So the
    observed difference is not a by-product of baseline differences; it is the genuine intervention effect. ANCOVA is
    the standard way to fairly control baseline differences in before/after landscape intervention studies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Sum-of-squares type (I/II/III).
    - Homogeneity-of-regression-slopes test (factor x covariate interaction — the key ANCOVA assumption).
    - Effect sizes: omega-squared, epsilon-squared.
    - Levene & residual Shapiro; Bonferroni post-hoc on adjusted (estimated marginal) means.

====================================================================================

#11  Bootstrap CI
    file: 11_bootstrap_ci_visitor_median.xlsx
  >> SCENARIO (narration):
    We want a confidence interval for the median of daily visitors; the data are
    skewed, so we use bootstrap (resampling), which makes no distributional assumption.
  >> VARIABLE SELECTION:
    - Variable: visitor_daily
    - Statistic: median

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed median daily visitors = 266
    95% Bootstrap CI (via resampling)

>> COMMENTARY (narration):
    We derived a confidence interval for the median of daily visitors by bootstrapping: the observed median is 266.
    Visitor data are typically right-skewed (most days moderate, some very busy), and classic normal-assuming formulas
    can be unreliable for such data. Bootstrap estimates the interval without distributional assumptions, through
    thousands of resamples. It is a robust way to express the uncertainty of central tendency in skewed use data
    (visitors, daily occupancy); the median is a more representative measure than the mean in a skewed distribution.

====================================================================================

#12  Permutation Test
    file: 12_permutation_green_area_stress.xlsx
  >> SCENARIO (narration):
    We test green-space exposure's effect on stress with a permutation test, which
    shuffles group labels thousands of times and makes no distributional assumption.
  >> VARIABLE SELECTION:
    - Grouping: group
    - Measure: stress

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    t = 2.89   p = 0.0057 **   H0 REJECTED
    Stress differs significantly between the two groups (high/low green exposure)

>> COMMENTARY (narration):
    We tested the effect of green-space exposure on stress with a permutation test. Permutation shuffles the group
    labels thousands of times to measure whether the observed difference could arise by chance -- requiring no
    distributional assumption. The result is significant (p = 0.006): the high- and low-green-exposure groups' stress
    levels differ more than chance allows. This is quantified evidence of landscape's "restorative" effect -- the
    stress-reducing role of green space (a core finding of therapeutic-landscape/restorative-environment research). In
    small samples permutation is a robust, assumption-free alternative to the t-test.

====================================================================================

#13  Multiple Comparison
    file: 13_multiple_comparison_plant_vitality.xlsx
  >> SCENARIO (narration):
    Comparing treatments pairwise on plant vitality, we control the inflated false-
    positive risk with Bonferroni/Holm/FDR corrections.
  >> VARIABLE SELECTION:
    - Grouping: treatment
    - Measure: vitality_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Plant-vitality-ratio groups vs reference: TOTAL SIGNIFICANT = 3/8 (in every correction method)

>> COMMENTARY (narration):
    When we compare different treatments' effect on plant vitality ratio pairwise, we run many tests -- each carrying
    a false-positive risk. Multiple-comparison correction controls that risk. Here three of eight comparisons stayed
    significant under every correction method (Bonferroni, Holm, FDR) -- 3/8. So some treatments genuinely differ from
    the reference, while others lose significance after correction. In planting/cultivation trials, skipping this step
    risks reporting spurious differences; the correct correction is essential for reliable treatment selection.

====================================================================================

#14  Mann-Whitney U
    file: 14_mann_whitney_city_suburb_aesthetic.xlsx
  >> SCENARIO (narration):
    We compare city-centre and suburb parks' aesthetic scores; aesthetic scores are
    ordinal and may not be normal, so the nonparametric Mann-Whitney is appropriate.
  >> VARIABLE SELECTION:
    - Grouping: location
    - Measure: aesthetic_score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Mann-Whitney U = 670.50   p = 0.041 *   r = .26
    Aesthetic score differs significantly by city/suburb location (borderline).   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether park aesthetic score differs by city-centre vs suburb location with the nonparametric Mann-
    Whitney -- because aesthetic scores are ordinal and may not be normal. The result is borderline significant (U =
    670.5, p = 0.041, r = 0.26): a difference exists but the p-value is close to 0.05 and the effect is moderate.
    City and suburb parks differ in aesthetic perception (probably design/maintenance/context). For ordinal/skewed
    aesthetic data Mann-Whitney is more reliable than the t-test and is widely used in landscape perception/preference
    studies. A borderline result should be confirmed with a larger sample.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Computation method: auto / exact (precise for small n) / asymptotic.
    - Effect sizes: CLES and Z/sqrt(N) (rank-biserial r already shown).
    - Descriptives (median etc.).

====================================================================================

#15  Wilcoxon Signed-Rank
    file: 15_wilcoxon_maintenance_cleanliness.xlsx
  >> SCENARIO (narration):
    We measure a maintenance program's effect by comparing the same parks' cleanliness
    before and after. For paired, ordinal/skewed data the paired Wilcoxon is appropriate.
  >> VARIABLE SELECTION:
    - Before: cleanliness_before
    - After: cleanliness_post

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Wilcoxon W = 0.00   p < .001 ***   r = 0.892   n (nonzero) = 30
    Cleanliness score changed significantly before vs after maintenance.   H0 REJECTED

>> COMMENTARY (narration):
    We measured a maintenance program's effect on park cleanliness by comparing the same parks' cleanliness scores
    before and after with the paired Wilcoxon test. The result is very strong: W = 0, p < .001, effect size r = 0.89.
    Cleanliness improved in the same direction and by a large amount in nearly all parks -- maintenance had a clear
    effect. For paired measures and ordinal/skewed score data, Wilcoxon is the right choice. It is a robust,
    assumption-free way to measure the effect of park-maintenance interventions; used in landscape management to
    evaluate the effectiveness of maintenance policies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Hypothesis direction; continuity correction.
    - Zero-difference handling: wilcox (drop) / pratt / zsplit.
    - Descriptives for both measures and their difference.

====================================================================================

#16  Kruskal-Wallis
    file: 16_kruskal_soil_survival.xlsx
  >> SCENARIO (narration):
    We examine whether plant survival duration varies by soil type; for more than two
    groups and possibly non-normal duration data, Kruskal-Wallis is appropriate.
  >> VARIABLE SELECTION:
    - Grouping: soil_type
    - Measure: survival_year

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    H(2) = 4.4582   p = 0.108 ns   eta-squared_H = 0.06
    Plant survival duration does NOT differ significantly by soil type.   H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether plant survival duration varies by soil type with Kruskal-Wallis -- the nonparametric ANOVA for
    more than two groups. The result is non-significant (H(2) = 4.46, p = 0.108): there is no significant difference
    in survival among soil types. A non-significant result is information too -- in this data plants show similar
    survival across soil types (perhaps correct species choice/planting technique compensates for soil differences).
    Because count/ordinal duration data may not be normally distributed, Kruskal-Wallis is more appropriate than
    classic ANOVA. Negative findings help avoid over-focusing on soil type when making planting decisions.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Epsilon-squared effect size.
    - Dunn post-hoc pairwise comparison (tie-corrected, Bonferroni/Holm).
    - Descriptives per group.

====================================================================================

#17  Friedman
    file: 17_friedman_panelist_composition.xlsx
  >> SCENARIO (narration):
    We compare the same panelists' aesthetic scores for four landscape compositions.
    Same raters score all compositions (repeated, ordinal), so the Friedman test is appropriate.
  >> VARIABLE SELECTION:
    - Repeated measures: composition_A_aesthetic, composition_B_aesthetic, composition_C_aesthetic, composition_D_aesthetic

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Friedman chi-square(3) = 41.91   p < .001 ***   Kendall W = 0.4657   n = 30
    The same panelists' aesthetic scores for 4 compositions differ significantly.   H0 REJECTED

>> COMMENTARY (narration):
    We compared the same panelists' aesthetic scores for four different landscape compositions with the Friedman test
    -- the nonparametric counterpart of repeated-measures ANOVA. The result is highly significant (chi-square(3) =
    41.91, p < .001), with Kendall W = 0.47 indicating moderate-high consistency of the ranking across panelists: they
    largely agree on which composition is more aesthetic. Because the same panelists rate all compositions,
    observations are dependent; Friedman accounts for this correctly. In landscape aesthetic evaluation (visual-
    preference surveys), it is the right test for comparing design alternatives on the same raters.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pairwise Wilcoxon signed-rank post-hoc (Bonferroni/Holm).
    - Descriptives per condition.

====================================================================================

#18  Binomial Test
    file: 18_binomial_seedling_survival.xlsx
  >> SCENARIO (narration):
    We test whether planted seedlings' survival rate equals chance (50%). One binary
    proportion against a fixed reference calls for the binomial test.
  >> VARIABLE SELECTION:
    - Binary variable: survival_status
    - Expected proportion: 0.50

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed survival proportion = 0.7362   test proportion = 0.50
    p < .001 ***   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the survival rate of planted seedlings equals chance (50%) with the binomial test. The observed
    proportion is 0.74 -- three-quarters of seedlings survived. This differs significantly from 50% (p < .001), so
    planting success is good. The binomial test is the simplest way to compare a single binary proportion (survived/
    died, present/absent) against a known reference. In landscape practice, comparing planting/seedling survival rates
    against a target or acceptable threshold is a practical way to evaluate afforestation/greening success.

====================================================================================

#19  Sign Test
    file: 19_sign_test_pruning_crown.xlsx
  >> SCENARIO (narration):
    We measure pruning's effect on crown diameter by comparing the same trees before
    and after. The assumption-free sign test, focused on the direction of change, is appropriate.
  >> VARIABLE SELECTION:
    - Before: crown_diameter_before_m
    - After: crown_diameter_post_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Positive (before > after) = 22   (crown diameter mostly DECREASED)   p = 0.0037 **   H0 REJECTED

>> COMMENTARY (narration):
    We measured a pruning intervention's effect on tree crown diameter by comparing the same trees' crown diameter
    before and after with the paired sign test. In most trees (22) the crown diameter decreased (before was larger),
    p = 0.004. So pruning reduced crown diameter as expected. The sign test looks only at the direction of change
    (decreased/increased), not its magnitude; it is a robust, assumption-free test. When the direction of change
    matters, or when the distribution is very skewed, it is a coarse but sturdy before/after indicator. It is used to
    measure the effect of tree care/pruning operations.

====================================================================================

#20  Runs Test
    file: 20_runs_test_park_occupancy.xlsx
  >> SCENARIO (narration):
    We test whether the daily park-occupancy series is random or shows clustering/
    pattern; the Wald-Wolfowitz runs test checks for a pattern in successive observations.
  >> VARIABLE SELECTION:
    - Series: visitor_count

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Observed runs = 32   Z = 0.27   (accepted as random)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We tested whether the daily park-occupancy (visitor) series is randomly distributed -- whether successive days
    cluster or show a pattern -- with the Wald-Wolfowitz runs test. With Z = 0.27 the observed number of runs is very
    close to expected; the series is accepted as random, with no systematic pattern. This matters methodologically:
    if a series is non-random (autocorrelated), many classic tests become invalid. Because randomness holds here,
    standard methods can be used safely. The runs test is ideal for this pre-check on time-ordered use data.

====================================================================================

#21  Chi-Square Independence
    file: 21_chisquare_type_disease.xlsx
  >> SCENARIO (narration):
    We investigate whether tree type is associated with disease occurrence. For
    association between two categorical variables, the chi-square test of independence is appropriate.
  >> VARIABLE SELECTION:
    - Row: type
    - Column: disease

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 49.7989   p < .001 ***   Cramer V = 0.39 (Strong)
    Tree type is associated with disease.   H0 REJECTED

>> COMMENTARY (narration):
    We tested whether tree type is associated with disease occurrence using chi-square. The result is strong:
    chi-square(3) = 49.80, p < .001, Cramer V = 0.39. So type and disease are not independent -- some types are
    markedly more disease-prone. This is critical for landscape plant selection: choosing disease-resistant types
    directly affects urban tree health and maintenance cost. For association between two categorical variables the
    chi-square test of independence is the standard method; Cramer V measures the strength. Knowing the type-disease
    pattern guides building a resistant species palette.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Yates continuity correction (2x2 tables).
    - G-test (likelihood-ratio chi-square) alternative.
    - Effect sizes: phi (2x2) and contingency coefficient C (besides Cramer's V).
    - Expected-counts table; standardized residuals (|>2| flags the deviating cell).

====================================================================================

#22  Chi-Square Goodness-of-Fit
    file: 22_chisquare_goodnessfit_type_distribution.xlsx
  >> SCENARIO (narration):
    We test whether plant types in an area are equally distributed; the chi-square
    goodness-of-fit test compares the observed category distribution to the equal expectation.
  >> VARIABLE SELECTION:
    - Categorical variable: type

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(3) = 14.88   p = 0.002 **   n = 200   k = 4 types (expected: equal 1/k)
    H0 REJECTED

>> COMMENTARY (narration):
    We tested whether the plant types in a park/city are equally distributed with the chi-square goodness-of-fit test.
    The observed distribution deviates significantly from the equal expectation (chi-square(3) = 14.88, p = 0.002):
    some types are more common than expected, others rare. So the type distribution is not balanced. This matters for
    species diversity (biodiversity) in landscape: dominance by a single type increases monoculture and disease/pest-
    outbreak risk. The goodness-of-fit test is the right way to compare an observed category distribution against a
    theoretical expectation (equal or target proportions); used to evaluate species-diversification policies.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Effect sizes: Cohen's w and Cramer's V.
    - G-test (likelihood ratio) alternative.
    - Standardized residuals per category (|>2| = notable deviation).

====================================================================================

#23  Fisher's Exact Test
    file: 23_fisher_seedling_method.xlsx
  >> SCENARIO (narration):
    We examine the association between planting method and retention in a small sample;
    chi-square is unreliable for small-cell 2x2 tables, so Fisher's exact test is appropriate.
  >> VARIABLE SELECTION:
    - Row: method
    - Column: retention

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Fisher Exact p = 0.0157 *   Odds Ratio = 0.073   Phi = 0.49 (Strong)
    Planting method is associated with retention.   H0 REJECTED

>> COMMENTARY (narration):
    We tested the association between seedling planting method and retention (whether the seedling takes) with
    Fisher's exact test -- because chi-square is unreliable for tables with few observations (small cells). The result
    is significant (p = 0.016), odds ratio 0.073: the chance of retention differs markedly between methods. Phi = 0.49
    indicates a strong association. This shows that in landscape practice the planting technique (containerised/bare-
    root, season, irrigation) strongly affects retention success. In small-sample planting pilot trials, Fisher's
    exact test gives more accurate results than chi-square.

====================================================================================

#24  McNemar Test
    file: 24_mcnemar_inventory_method.xlsx
  >> SCENARIO (narration):
    We compare an old and a new inventory method on the same areas. For paired binary
    data, McNemar's test (before/after change) is appropriate.
  >> VARIABLE SELECTION:
    - Before: old_method
    - After: new_method

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    McNemar (exact binomial) min(b,c) = 5   p = 0.0266 *   discordant: b = 16, c = 5   Cohen g = 0.52
    H0 REJECTED

>> COMMENTARY (narration):
    We compared an old and a new inventory (area-detection) method on the same areas -- paired binary data call for
    McNemar's test. The result is significant (p = 0.027): the new method caught what the old missed in 16 areas,
    while the old caught the new in only 5 -- a clear advantage for the new method. So the new method (e.g. satellite/
    drone-based detection) is more sensitive than the old. McNemar looks only at the cells that changed (discordant
    cells); it is the right way to compare two methods on the same areas. It is used to prove the superiority of new
    technologies (remote sensing) over old ones in landscape inventory.

====================================================================================

#25  Cohen's Kappa
    file: 25_kappa_maintenance_expert.xlsx
  >> SCENARIO (narration):
    We measure the agreement of two experts independently assessing the same parks'
    maintenance with the chance-corrected Cohen's Kappa.
  >> VARIABLE SELECTION:
    - Rater 1: expert_a
    - Rater 2: expert_b

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's Kappa = 0.6771 (substantial/strong agreement)   p < .001 ***
    Two experts' park-maintenance assessment

>> COMMENTARY (narration):
    We measured the agreement between two experts independently assessing the same parks' maintenance with Cohen's
    Kappa. Kappa = 0.68 is in the "substantial/strong" range -- a chance-corrected measure of agreement, i.e. the real
    consistency after removing accidental agreement. In landscape management this matters: if maintenance assessment
    varied from expert to expert, maintenance priorities and budget decisions would be unreliable. Kappa = 0.68 shows
    the assessments are largely consistent. Reporting inter-rater reliability is the quality assurance of park quality/
    maintenance audits.

====================================================================================

#26  Cochran-Mantel-Haenszel
    file: 26_cmh_maintenance_disease_region.xlsx
  >> SCENARIO (narration):
    We examine the maintenance-disease association while controlling for region (strata);
    the CMH test stratifies by a third variable.
  >> VARIABLE SELECTION:
    - Row: maintenance_density
    - Column: disease
    - Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CMH chi-square(1) = 20.69   p < .001 ***   Common OR (MH) = 3.13  (95% CI: 1.89 - 5.19)
    Breslow-Day chi-square(2) = 2.84, p = 0.24 (homogeneous OR)   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between maintenance intensity and plant disease while controlling for regions (strata)
    with the CMH test. Even holding region constant the association is strong: common odds ratio 3.13 (p < .001) --
    stripped of the region confounder, low-maintenance parks have 3 times the odds of disease as high-maintenance ones.
    The Breslow-Day test is non-significant (p = 0.24), meaning the association is consistent across all regions
    (homogeneous OR). This proves maintenance's protective role for plant health, region-independently. CMH is the
    classic way to control for a third variable (region) by stratification.

====================================================================================

#27  Log-Linear Analysis
    file: 27_log_linear_design_region_user.xlsx
  >> SCENARIO (narration):
    We analyse the multi-way categorical relationships and interactions among design,
    region and user type; log-linear analysis suits more than two categorical variables.
  >> VARIABLE SELECTION:
    - Factors: design, region, user_type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Best model AIC = 75.90   Pearson chi-square = 0.43 (model fits the data very well)
    Factors: design x region x user type

>> COMMENTARY (narration):
    We analysed the contingency structure formed jointly by three categorical variables -- design, region and user
    type -- with a log-linear model. The selected model fits very well (Pearson chi-square = 0.43, low), with AIC =
    75.90 as the most parsimonious model. Log-linear analysis lets us untangle interactions among more than two
    categorical variables (which are dependent, which interaction is significant) -- it is like ANOVA for categorical
    data. In landscape it is powerful for summarising the multi-way relationships among design type, region and user
    profile; used in user-centred design and target-audience analysis.

====================================================================================

#28  Cross-Tabulation
    file: 28_cross_age_purpose.xlsx
  >> SCENARIO (narration):
    We want to see the co-occurrence of age group and visit purpose; a crosstab and
    chi-square examine the association of two categorical variables.
  >> VARIABLE SELECTION:
    - Row: age_group
    - Column: visit_purpose

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    chi-square(6) = 54.14   p < .001 ***   Cramer V = 0.28 (Moderate)
    Age group is associated with visit purpose.   H0 REJECTED

>> COMMENTARY (narration):
    We examined the association between age group and visit purpose (sport, rest, socialising, children's play, etc.)
    with a crosstab and chi-square. The association is significant and moderate (chi-square(6) = 54.14, p < .001,
    Cramer V = 0.28): age groups use parks for different purposes -- e.g. youth for sport/socialising, the elderly for
    rest, families for children's play. This is a critical finding for age-sensitive (all-ages) space design in
    landscape. The crosstab is the most basic and readable way to see the co-occurrence of two categorical variables;
    Cramer V gives the strength.

====================================================================================

#29  Multiple Response - Frequency
    file: 29_mr_frequency_activities.xlsx
  >> SCENARIO (narration):
    We summarise as frequencies the activities users do in the park (a multiple-choice
    question where several options can be ticked).
  >> VARIABLE SELECTION:
    - Multiple-response column: activities

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 250   Respondents = 250 (100%)
    Park activities (walking, picnic, sport, children's play, dog walking...) selected as multi-choice

>> COMMENTARY (narration):
    Park users were asked which activities they do in the park -- a "multiple response" question where several options
    can be ticked. All 250 users selected at least one activity, so the response rate is 100%. Multiple-response
    frequency analysis shows how many times and by what percentage of users each activity was chosen; the totals
    exceed 100% because everyone can tick several boxes. Profiling users' activity use is very useful in landscape
    planning to determine which activities need dedicated equipment/spaces.

====================================================================================

#30  Multiple Response - Crosstab
    file: 30_mr_categorical_activity_region.xlsx
  >> SCENARIO (narration):
    We cross-tabulate multiple-choice activities by the park's region to see which
    region prefers which activities.
  >> VARIABLE SELECTION:
    - Multiple-response column: activities
    - Grouping: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 220   Group column: region
    Park activities crossed by region

>> COMMENTARY (narration):
    We cross-tabulated park activities (multiple-choice) by the park's region. With data from 220 users we can see how
    each activity is distributed across regions. The multiple-response crosstab answers "which region prefers which
    activities more" -- e.g. sport leads in one region, picnic in another. This guides region-specific (place-
    sensitive) park programming: different equipment/space prioritisation per region according to its use pattern.

====================================================================================

#31  Multiple Response x Multiple Response
    file: 31_mr_mr_activity_equipment.xlsx
  >> SCENARIO (narration):
    We cross two multiple-choice questions (activities done and equipment brought) to
    examine activity-equipment pairings.
  >> VARIABLE SELECTION:
    - Multiple response 1: activities
    - Multiple response 2: equipment

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total cases n = 180
    Activities x equipment (backpack, camera, bicycle, tent) cross-tabulated

>> COMMENTARY (narration):
    This is the most advanced multiple-response analysis: we cross two separate multi-choice questions against each
    other -- the activities the user does and the equipment they bring. Among 180 users we can see which activity-doers
    bring which equipment: e.g. those bringing a bicycle do cycling, those bringing a tent/backpack do picnic/camp.
    This table reveals activity-equipment pairings, providing powerful data for park-equipment planning (bicycle
    parking, picnic areas, camera/photo spots).

====================================================================================

#32  Cochran's Q
    file: 32_cochran_q_safety_site.xlsx
  >> SCENARIO (narration):
    We compare the same visitors' feeling of safety (binary) across different park
    sites; for repeated binary measures Cochran's Q is appropriate.
  >> VARIABLE SELECTION:
    - Subject: visitor_id
    - Condition (site): site
    - Binary outcome: safe_feels

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cochran's Q = 0.00   p = 1.000   (NO difference in safety feeling across sites)
    H0 NOT REJECTED

>> COMMENTARY (narration):
    We compared the same visitors' feeling of safety (yes/no) across different park sites with Cochran's Q test -- the
    multi-occasion counterpart of repeated binary measures. The result is non-significant (Q = 0, p = 1.0): there is
    no difference in safety feeling across sites -- visitors perceive similar safety everywhere. A non-significant
    result is valuable too; it says the sites do not differ in safety perception in this data (perhaps consistent
    safety design/lighting). Cochran's Q is the right way to test change in repeated binary (yes/no) measurements on
    the same people.

====================================================================================

#33  Correlation Matrix
    file: 33_correlation_park_5degisken.xlsx
  >> SCENARIO (narration):
    We want to see the relationships and collinearity among five park variables (area,
    green ratio, visitors, noise, satisfaction) before modelling; a correlation matrix is appropriate.
  >> VARIABLE SELECTION:
    - Variables: area_m2, green_area_ratio, visitor_weekly, noise_dB, satisfaction

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Strongest relationship: area - weekly visitors   r = +0.918   p < .001 ***  (Strong) -- 3 strong pairs
    5 variables: area, green ratio, visitors, noise, satisfaction

>> COMMENTARY (narration):
    We computed all pairwise correlations among five park variables. The strongest is between area and weekly visitors
    (r = 0.92) -- larger parks attract more visitors, an expected result. Three pairs exceed the |r| >= 0.75
    threshold, so some variables are tightly linked (area-visitors-green). This warns us to watch for collinearity
    when building a regression. The correlation matrix is the first step before modelling to see the structure among
    variables. In landscape, understanding the relationships among area, green fabric, use and satisfaction is the
    basis of models predicting park performance.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - p-values are now produced for ALL methods (Pearson/Spearman/Kendall), not only Pearson.
    - Hypothesis direction (two-sided / right / left).
    - Multiple-comparison p-adjustment across pairs: Bonferroni / Holm / FDR (Benjamini-Hochberg).
    - Confidence interval for r via Fisher z (Pearson/Spearman) or Kendall SE.

====================================================================================

#34  Bland-Altman
    file: 34_bland_altman_green_measurement.xlsx
  >> SCENARIO (narration):
    We assess the agreement of two methods of measuring green-area ratio (satellite vs
    hand measurement); Bland-Altman, based on the difference between measurements, is appropriate.
  >> VARIABLE SELECTION:
    - Method 1: green_satellite
    - Method 2: green_hand_measurement

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Satellite mean = 0.518   |   Hand measurement mean = 0.512   |   Bias ~ +0.006 (methods agree very well)

>> COMMENTARY (narration):
    We compared two methods of measuring green-area ratio -- satellite/remote sensing vs hand (field) measurement --
    with Bland-Altman. This method does not look at correlation (high correlation does not guarantee agreement); it
    measures the difference between the two measurements. The means are almost identical (0.518 vs 0.512), the bias
    just 0.006 -- satellite and hand measurement give almost the same result and are practically interchangeable. This
    is very valuable: satellite-based measurement can safely replace expensive, time-consuming field measurement.
    Bland-Altman is the standard way to assess the agreement of two measurement methods (remote sensing vs field);
    it is indispensable in landscape/GIS method validation.

====================================================================================

#35  Effect Size (Cohen's d)
    file: 35_effect_size_irrigation_growth.xlsx
  >> SCENARIO (narration):
    We want not just the significance but the practical magnitude (effect size) of the
    growth difference between two irrigation methods.
  >> VARIABLE SELECTION:
    - Measure: growth_cm
    - Grouping: irrigation_method

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cohen's d = -0.94  (large effect)
    Growth (cm) difference between two irrigation methods

>> COMMENTARY (narration):
    We measured the plant-growth difference between two irrigation methods (e.g. drip vs surface) not just as "is it
    significant" but "how large is it" -- as an effect size. Cohen's d = -0.94 is a large effect; the two methods
    differ markedly in growth. P-values inflate with sample size, but effect size is sample-independent and gives the
    practical importance of the real difference. In landscape, "how much difference does this irrigation method make"
    is critical for balancing water saving and plant health; a large effect of d = 0.94 says the irrigation method
    makes a notable difference. Reporting effect size is standard in publications.

====================================================================================

#36  Canonical Correlation (CCA)
    file: 36_cca_physical_perception.xlsx
  >> SCENARIO (narration):
    We examine the multivariate relationship between the park's physical-feature set and
    the user-perception set; for the relationship between two blocks of variables, CCA is appropriate.
  >> VARIABLE SELECTION:
    - X set (physical): green_ratio, water_area_pct, tree_density, path_length_m, ruggedness_index, relief
    - Y set (perception): aesthetic, safety, repeat_visit

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CC1: r = 0.983, chi-square(18) = 1199.24, p < .001   |   CC2: r = 0.967, p < .001
    X set: physical (green/water/tree-density/path-length/ruggedness/relief)  |  Y set: perception (aesthetic/safety/repeat-visit)

>> COMMENTARY (narration):
    We related two multivariate sets -- the park's physical features (green ratio, water area, tree density, path
    length, ruggedness, relief) and user perception (aesthetic, safety, repeat visit) -- with canonical correlation.
    The first canonical axis is extremely strong (r = 0.98, p < .001), the second also significant (r = 0.97): physical
    design and user perception are very tightly linked. So a park's physical features strongly determine the user's
    aesthetic/safety/loyalty perception. CCA is the most direct answer to "how do two blocks of variables co-vary"; in
    landscape it is a powerful method for studying "how design features shape perception" -- the central question of
    evidence-based design.

====================================================================================

#37  Correspondence Analysis
    file: 37_ca_park_purpose.xlsx
  >> SCENARIO (narration):
    We want to visualise the contingency relationship between park type and visit
    purpose on a two-dimensional map; correspondence analysis is appropriate.
  >> VARIABLE SELECTION:
    - Row: park_type
    - Column: visit_purpose

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total inertia = 1.293   Number of dimensions = 2
    Park type x visit purpose contingency table

>> COMMENTARY (narration):
    We projected the contingency table of park type by visit purpose into a two-dimensional map with correspondence
    analysis. Total inertia 1.293 indicates a very strong association; the two dimensions visualise most of that
    structure. Type-purpose pairs that fall close on the map indicate that the park type is associated with that
    purpose -- e.g. neighbourhood park with children's play/rest, city park with sport/events, coastal park with
    view/walking. This lets us read at a glance how landscape typology overlaps with use purposes. Correspondence
    analysis is a powerful way to visualise categorical relationships and gives the association strength (inertia)
    numerically.

====================================================================================

#38  Variable Clustering (VarClus)
    file: 38_varclus_18_item.xlsx
  >> SCENARIO (narration):
    We cluster the 18 park-quality items by which items measure the same dimension;
    VarClus clusters variables (not observations).
  >> VARIABLE SELECTION:
    - Variables: aesthetic_m1..m6, function_m1..m6, environment_m1..m6 (18 items)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 3   (18 items: aesthetic / function / environment)
    A representative item was selected from each cluster

>> COMMENTARY (narration):
    We clustered the 18 items measuring park quality by similarity into three groups. VarClus clusters items, not
    observations: it shows which items measure the same "dimension". The three clusters formed exactly as expected:
    aesthetic, functional and environmental quality items clustered separately. This confirms the quality scale really
    measures three sub-dimensions and lets us pick one representative per group to reduce the number of items. In
    developing landscape quality scales and in item reduction, it is extremely practical for weeding out redundancy.

====================================================================================

#39  Multiple Linear Regression
    file: 39_multiple_regression_visitor.xlsx
  >> SCENARIO (narration):
    We predict weekly visitors from several design/management variables; multiple
    linear regression explains a continuous outcome with several predictors.
  >> VARIABLE SELECTION:
    - Target: visitor_weekly
    - Predictors: area_m2, green_area_ratio, access_score, maintenance_score

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.962   Adjusted R2 = 0.961
    Weekly visitors ~ area + green ratio + access score + maintenance score

>> COMMENTARY (narration):
    We built a multiple linear regression predicting weekly visitors from four variables -- area, green ratio, access
    score, maintenance score. The model is extremely strong: together they explain 96% of visitor variance (R2 =
    0.96). This is a strikingly high explanatory rate -- park use is largely determined by these measurable design/
    management variables. Such a model shows which factor attracts the most visitors and lets us test "what happens to
    visitors if we improve access/maintenance by so much" scenarios. Multiple regression is the basic tool for
    explaining park performance (use) with design and management drivers.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Robust standard errors (HC0-HC3): heteroscedasticity-robust SE; HC3 recommended for small n.
    - Standardized (beta) coefficients to compare relative effect.
    - (Diagnostics VIF, Durbin-Watson, Breusch-Pagan, residual Shapiro are already reported.)

====================================================================================

#40  Logistic Regression
    file: 40_logistic_repeat_visit.xlsx
  >> SCENARIO (narration):
    We predict whether a user revisits the park (binary) from satisfaction, stay
    duration and age; logistic regression suits a binary outcome.
  >> VARIABLE SELECTION:
    - Target (binary): repeat_visit
    - Predictors: satisfaction, stay_min, age

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: Logistic Regression   Pseudo R2 = 0.05
    Repeat visit (1/0) ~ satisfaction + stay duration + age

>> COMMENTARY (narration):
    We built a logistic regression predicting whether a user will revisit the park from satisfaction, stay duration
    and age. Because the outcome is binary (revisited/not), logistic regression is the right choice. Pseudo R2 = 0.05
    is low -- these variables weakly explain repeat visits; so repeat visits may be determined by other (unmeasured)
    factors (distance, alternative parks, habit). A low pseudo R2 is information too: it says "the variables in the
    model are insufficient". In landscape, predicting user loyalty (repeat visits) matters for understanding the
    factors that build park loyalty; logistic regression is the right method for binary outcomes.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Pseudo-R squared: Cox-Snell and Nagelkerke (besides McFadden).
    - Classification metrics: accuracy / sensitivity / specificity / AUC (cutoff 0.5).
    - Hosmer-Lemeshow goodness-of-fit test; VIF for predictors.

====================================================================================

#41  Count (Poisson) Regression
    file: 41_poisson_bird_sayim.xlsx
  >> SCENARIO (narration):
    We predict bird count (a count variable) from green ratio and water-area presence;
    since count data are not normal, Poisson regression is appropriate.
  >> VARIABLE SELECTION:
    - Count (target): bird_count
    - Predictors: green_ratio, water_area_present

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 884.69   Deviance = 185.78
    Bird count ~ green ratio + water-area presence   (Poisson)

>> COMMENTARY (narration):
    We built a Poisson regression predicting bird count in a park -- a count variable -- from green ratio and water-
    area presence. Count data (0,1,2,...) are not normally distributed, so we use Poisson rather than linear
    regression. The model typically shows bird count rising with green ratio and water area. This is quantified
    evidence of landscape's ecosystem-service (biodiversity support) dimension -- green, watered parks host more birds.
    Count regression is the correct way to relate biodiversity counts (birds, butterflies, species richness) to design
    variables; it is used in urban ecology and biophilic-design studies.

====================================================================================

#42  Multinomial Logistic Regression
    file: 42_multinomial_park_preference.xlsx
  >> SCENARIO (narration):
    We predict users' park-type choice (several unordered categories) from age and
    green preference; multinomial logistic regression is appropriate.
  >> VARIABLE SELECTION:
    - Target (categorical): park_type_choice
    - Predictors: age, green_preference

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 385.24   Reference class: 'modern'
    Park-type choice ~ age + green preference

>> COMMENTARY (narration):
    We modelled users' park-type choice (several categories) from age and green preference with multinomial logistic
    regression. Because there are more than two unordered categories, this method fits: 'modern' is the reference, and
    a separate equation is built for each other type. Coefficients read as "how the odds of choosing a given park type
    change as age/green-preference shift". This shows which user profiles prefer which park type -- a valuable tool for
    target-audience-focused park design and diversified park-portfolio planning.

====================================================================================

#43  Ordinal Logistic Regression
    file: 43_ordinal_maintenance_quality.xlsx
  >> SCENARIO (narration):
    We predict maintenance quality level (low<medium<high, ordered) from annual budget
    and weekly maintenance days; ordinal logistic suits an ordered outcome.
  >> VARIABLE SELECTION:
    - Target (ordered): maintenance_quality_level
    - Predictors: annual_budget_TL, weekly_maintenance_day

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 334.58
    Maintenance quality level (low<medium<high) ~ annual budget + weekly maintenance days

>> COMMENTARY (narration):
    We built an ordinal logistic regression predicting maintenance quality level (low, medium, high -- ordered) from
    annual budget and weekly maintenance days. These classes are ordered, and the ordinal model uses that order,
    making it more powerful and interpretable than multinomial. The coefficients give the tendency to move to a higher
    quality level as budget and maintenance days increase (resource-quality relationship). For ordered categorical
    outcomes (quality tier, satisfaction level), ordinal logistic is the right method; in landscape management it is
    used to model the maintenance-budget-quality relationship.

====================================================================================

#44  PLS Regression
    file: 44_pls_spectral_NDVI.xlsx
  >> SCENARIO (narration):
    We predict the vegetation index (NDVI) from 12 correlated spectral bands; PLS
    regression suits collinear predictors.
  >> VARIABLE SELECTION:
    - Target: NDVI_dense
    - Predictors: band_01..band_12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2_train = 0.918   R2_CV (5-fold) = 0.912
    NDVI (vegetation index) ~ 12 spectral bands

>> COMMENTARY (narration):
    We predicted vegetation density (NDVI) from 12 spectral satellite bands with PLS (Partial Least Squares)
    regression. Spectral bands are highly correlated, so ordinary regression becomes unstable; PLS solves this by
    reducing the bands to a few orthogonal components. The model is very strong both in training (R2 = 0.92) and
    cross-validation (R2_CV = 0.91) -- no overfitting, excellent generalisation. Remote-sensing/GIS data are typically
    high-dimensional and collinear (neighbouring bands are similar); PLS keeps both predictive power and
    interpretability for such data. It is used in landscape for vegetation-cover and green-area mapping.

====================================================================================

#45  Probit Regression
    file: 45_probit_herbisit_dose_response.xlsx
  >> SCENARIO (narration):
    We model whether a herbicide dose kills weeds (binary dose-response); probit
    regression, assuming a normal-distribution curve, is appropriate.
  >> VARIABLE SELECTION:
    - Target (binary): death
    - Predictor: dose_lt_ha

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 164.67   Pseudo R2 (McFadden) = 0.349
    Death (1/0) ~ herbicide dose (lt/ha)

>> COMMENTARY (narration):
    We modelled whether a herbicide dose kills weeds with probit regression. Probit models a binary outcome like
    logistic but assumes a normal-distribution curve (cumulative normal) -- the classic model for toxicology/herbicide
    dose-response studies. Pseudo R2 = 0.349 is strong: dose largely explains the probability of death -- a classic
    dose-response relationship. The probit curve helps find the dose at which the death probability reaches 50% (LD50).
    In landscape/green-space management, determining the optimal (effective yet minimal) herbicide dose -- for both
    weed control and environmental-impact balance -- uses this model.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Classification metrics added: accuracy / sensitivity / specificity / AUC (marginal effects + McFadden already shown).

====================================================================================

#46  Tobit Regression
    file: 46_tobit_renewal_expenditure.xlsx
  >> SCENARIO (narration):
    We model renewal expenditure; because many parks have zero/floored (censored)
    expenditure, Tobit regression is appropriate.
  >> VARIABLE SELECTION:
    - Target (censored): renewal_expenditure_TL
    - Predictors: allocation_budget, renewal_importance

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AIC = 3444.11   Tobit (left-censored: expenditure floored)
    Renewal expenditure ~ allocation budget + renewal importance

>> COMMENTARY (narration):
    We modelled parks' renewal expenditure from allocation budget and a renewal-priority score. The problem: some
    parks get no renewal expenditure (zero) or expenditure piles up below a floor value -- this is censored data, and
    ordinary regression gives biased results. Tobit regression correctly models this left-censored structure (AIC =
    3444). Typically expenditure rises with budget and priority. The statistically correct solution for working with
    landscape/municipal budget data -- where many units have zero/low expenditure pile-up -- is the Tobit model.

====================================================================================

#47  Bayesian Linear Regression
    file: 47_bayesian_fertilizer_yield.xlsx
  >> SCENARIO (narration):
    We predict plant yield from fertiliser doses (N, P, K) while expressing uncertainty
    with a posterior distribution; Bayesian linear regression is appropriate.
  >> VARIABLE SELECTION:
    - Target: plant_yield_kg
    - Predictors: N_kg_ha, P_kg_ha, K_kg_ha

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Conjugate Normal-Inverse-Gamma prior   posterior sigma2 = 28.29
    Plant yield (kg) ~ nitrogen (N) + phosphorus (P) + potassium (K)

>> COMMENTARY (narration):
    We predicted plant yield from fertiliser doses (N, P, K) with Bayesian linear regression. The Bayesian advantage:
    the result is not a single p-value but a probability distribution (posterior) for the parameters -- we get a
    credible interval for each nutrient's effect and a "probability the effect is positive". The posterior error
    variance is sigma2 = 28.29. The Bayesian method honestly expresses uncertainty and can incorporate prior
    agricultural knowledge. In landscape/ornamental-plant cultivation, determining the optimal fertilisation regime,
    especially in low-replication settings, the Bayesian framework is valuable.

====================================================================================

#48  Nonlinear Regression (Gompertz)
    file: 48_nonlinear_gompertz_growth.xlsx
  >> SCENARIO (narration):
    Tree height-by-age is not linear but an S-curve saturating toward a ceiling; we use
    the Gompertz growth model.
  >> VARIABLE SELECTION:
    - Target: height_m
    - Predictor: age_year
    - Function: gompertz

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Model: y = a · exp(-b·exp(-c·x))  (Gompertz growth)   R2 = 0.989
    Tree height ~ age

>> COMMENTARY (narration):
    We modelled how tree height changes with age with the Gompertz growth model. Tree growth is not linear: it is slow
    when young, then fast, finally saturating toward a ceiling (a -- the asymptotic maximum height) -- an asymmetric
    S-curve. Gompertz captures exactly this form; the fit is excellent (R2 = 0.99). The curve gives three parameters:
    the ceiling (a), growth rate (c) and inflection (b). These are compared across species/conditions to answer "which
    tree grows faster/taller". Nonlinear regression is the basic tool of growth-age modelling in landscape/forest
    mensuration; a straight line would misrepresent the nature of tree growth.

====================================================================================

#49  Ridge Regression
    file: 49_ridge_soil_vitality.xlsx
  >> SCENARIO (narration):
    We predict plant vitality from 15 correlated soil parameters using Ridge regression,
    which is robust to collinearity.
  >> VARIABLE SELECTION:
    - Target: plant_vitality
    - Predictors: soil_p01..soil_p15

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Ridge (alpha = 1.0): R2 = 0.576   Adj. R2 = 0.542   n = 200
    Plant vitality ~ 15 soil parameters (mutually correlated)

>> COMMENTARY (narration):
    Predicting plant vitality from 15 highly correlated soil parameters (pH, N, P, K, organic matter, moisture, etc.),
    ordinary regression becomes unstable (collinearity inflates coefficients -- soil properties are naturally
    correlated). Ridge regression solves this by shrinking all coefficients -- it does not zero them but limits their
    magnitude, keeping prediction stable (R2 = 0.58). Ridge is ideal when "I want to keep all soil parameters but there
    is collinearity". For correlated soil/environmental measurement sets, it prevents overfitting and provides stable
    estimates.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (RidgeCV): selects the optimal regularization strength automatically.

====================================================================================

#50  Lasso Regression
    file: 50_lasso_30_pred_5_active.xlsx
  >> SCENARIO (narration):
    We want to select the few factors that truly determine design quality from 30
    candidates; Lasso performs automatic variable selection.
  >> VARIABLE SELECTION:
    - Target: design_quality
    - Predictors: x01..x30

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Lasso (alpha = 0.1): R2 = 0.645   Adj. R2 = 0.596   n = 250
    20 of 30 predictor coefficients shrunk to zero (automatic feature selection)

>> COMMENTARY (narration):
    In this data a few real predictors were mixed with many irrelevant variables. The strength of Lasso shows here:
    the penalty term drives irrelevant variables' coefficients exactly to zero -- automatic variable selection. 20 of
    30 predictors were eliminated, and the model stayed strong at R2 = 0.65. Lasso is extremely useful when you want
    to pick the few factors that truly determine design quality from many candidates -- in high-dimensional landscape/
    environmental data. While Ridge shrinks all variables, Lasso eliminates the irrelevant ones entirely.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Auto-alpha via cross-validation (LassoCV): selects the optimal regularization strength automatically.

====================================================================================

#51  Mediation Analysis
    file: 51_mediation_green_activity_satisfaction.xlsx
  >> SCENARIO (narration):
    We test whether green ratio's effect on satisfaction passes through physical
    activity (mediation).
  >> VARIABLE SELECTION:
    - Independent (X): green_ratio
    - Mediator (M): weekly_activity_min
    - Dependent (Y): satisfaction

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Indirect effect = 0.19   p = 0.303   95% CI (-0.17 , 0.57)   NOT SIGNIFICANT
    Path: green ratio -> weekly activity minutes -> satisfaction

>> COMMENTARY (narration):
    We tested whether green ratio's effect on satisfaction passes through physical activity with mediation analysis.
    This time the indirect effect is non-significant (0.19, p = 0.303, CI includes zero): green ratio's effect on
    satisfaction does not pass "through" physical activity. A non-significant mediation is information too -- the
    proposed mechanism (green space -> more activity -> more satisfaction) is not supported in this data; if green
    ratio affects satisfaction, it may go by another route (direct aesthetic/psychological effect). Mediation analysis
    can both confirm and refute a mechanism; a negative result prompts us to reconsider the assumed causal chain.

====================================================================================

#52  Path Analysis
    file: 52_path_landscape_quality.xlsx
  >> SCENARIO (narration):
    We test how aesthetic perception, use frequency and satisfaction affect repeat-visit
    intention in a single path model.
  >> VARIABLE SELECTION:
    - Model: repeat_visit_intention ~ aesthetic_perception + use_frequent + satisfaction ; satisfaction ~ aesthetic_perception

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 0.933   RMSEA = 0.213   (poor fit -- model should be revised)
    Model: aesthetic perception/use frequency/satisfaction -> repeat-visit intention

>> COMMENTARY (narration):
    Path analysis is extended regression that models several cause-effect relationships at once. Here we tested that
    aesthetic perception, use frequency and satisfaction affect repeat-visit intention, and that aesthetics feeds
    satisfaction. CFI = 0.93 is acceptable but RMSEA = 0.21 is high -- the hypothesised model does not fully match the
    data; some paths may be missing or extra. This is informative too: it says the theoretical model should be
    revised. Path analysis is a powerful way to separate direct and indirect effects and test landscape-quality/user-
    behaviour theories; both good and poor fit guide us in improving the model.

====================================================================================

#53  Linear Mixed Model (LMM)
    file: 53_lmm_park_year_green.xlsx
  >> SCENARIO (narration):
    Modelling the same parks' green ratio over years, we account for the dependence of
    within-park repeated measures with park as a random effect (LMM).
  >> VARIABLE SELECTION:
    - Dependent: green_ratio
    - Fixed effect: year
    - Random effect (group): park_id

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group variance (random intercept) = 0.0074   ICC = 0.9555
    Green ratio ~ year  +  (1 | park)

>> COMMENTARY (narration):
    Modelling the same parks' green ratio across years, the within-park repeated measures are not independent. The
    linear mixed model solves this by adding park as a random effect. ICC = 0.96 is very high: 96% of green-ratio
    variance comes from between-park differences; within-year change is very small. So parks are consistent within
    themselves (green ratio barely changes year to year), and the main difference is between parks. LMM is the correct,
    indispensable method for nested/repeated-measures data (park>year) -- it models each park's own trajectory and
    resolves the independence violation. It is used to track the stability of park green fabric over time.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Nakagawa marginal R-squared (fixed effects) and conditional R-squared (fixed + random), beside ICC.

====================================================================================

#54  Multiple Imputation
    file: 54_multiple_imputation_missing_inventory.xlsx
  >> SCENARIO (narration):
    We complete missing satisfaction values in the park inventory from other variables
    via multiple imputation (MAR assumption).
  >> VARIABLE SELECTION:
    - Dependent: satisfaction
    - Auxiliary variables: area_m2, green_ratio, visitor_weekly, noise_dB

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    m (number of imputations) = 5
    Missing inventory values estimated from area/green/visitors/noise

>> COMMENTARY (narration):
    In park-inventory data, missing observations are inevitable (unmeasured area, recording error); deleting them
    loses parks and biases results. Multiple imputation estimates the missing satisfaction values from other variables,
    creating 5 separate completed datasets, runs the analysis in each and combines the results -- thereby also
    accounting for estimation uncertainty. Under the MAR (missing at random) assumption this is the modern missing-data
    standard. When working with large park inventories, it is the soundest way to cope with lost records.

====================================================================================

#55  GEE
    file: 55_gee_park_visit_satisfaction.xlsx
  >> SCENARIO (narration):
    For repeated satisfaction measures from the same parks, we estimate the population-
    average relationship while correcting for clustering; GEE is appropriate.
  >> VARIABLE SELECTION:
    - Dependent: satisfaction
    - Predictor: visit
    - Group: park_id  |  Family: gaussian

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    visit coefficient = 0.076   p < .001 ***   QIC = 306.49
    Satisfaction ~ visit  (repeated measures, clustered within park)

>> COMMENTARY (narration):
    We took repeated (multi-visit) satisfaction measurements from the same parks -- one park's measurements are
    dependent. GEE estimates the relationship at the "population average" level in such clustered data, accounting for
    the correlation structure. The visit coefficient is significant and positive (0.076, p < .001): satisfaction rises
    with successive visits -- perhaps as maintenance/improvements progress or users get used to the park. GEE is an
    alternative to LMM: rather than modelling random effects, it corrects the correlation among observations and gives
    the marginal (average) effect. It is a common method for analysing repeated measures in longitudinal park-
    satisfaction monitoring.

====================================================================================

#56  GLMM
    file: 56_glmm_site_time_insect.xlsx
  >> SCENARIO (narration):
    We model insect counts (Poisson) with time/maintenance fixed effects and a site
    random effect; a mixed model + count distribution (GLMM) is appropriate.
  >> VARIABLE SELECTION:
    - Dependent (count): insect_count
    - Fixed effects: time, maintenance_present
    - Random effect (group): site_id  |  Family: poisson

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    maintenance_present coefficient = -0.462, p < .001 ***
    Poisson GLMM: insect count ~ time + maintenance presence  +  (1 | site)

>> COMMENTARY (narration):
    We modelled insect/pest counts (Poisson) with both time/maintenance fixed effects and a site random effect -- a
    GLMM: mixed model + count distribution. The maintenance effect is significant and negative (coef -0.46, p < .001):
    maintained sites have markedly fewer insects -- evidence of regular maintenance's role in pest control. GLMM
    combines the strengths of LMM (which assumes normality) and GLM (which ignores site clustering) for "repeated/
    nested count data". It is the correct analytic framework for landscape-health monitoring data like insect/pest
    counts nested within sites.

====================================================================================

#57  Elastic Net Regression
    file: 57_elasticnet_park_value.xlsx
  >> SCENARIO (narration):
    Predicting park value from 40 (partly correlated) predictors, we use Elastic Net for
    both selection and stability.
  >> VARIABLE SELECTION:
    - Target: park_value
    - Predictors: x01..x40  |  Method: elasticnet

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    alpha = 0.096   R2 (train) = 0.520
    Park value ~ 40 predictors

>> COMMENTARY (narration):
    Elastic Net is a hybrid of Ridge and Lasso: it both shrinks coefficients (like Ridge) and zeros out the irrelevant
    ones (like Lasso). Here we predicted park value from 40 predictors; Elastic Net keeps correlated variable groups
    together while eliminating the irrelevant ones. The model gives R2 = 0.52. When there are many collinear predictors
    -- area, location, environmental features, access variables -- Elastic Net is a balanced choice offering both
    selection and stability. It softens Lasso's "pick one at random from a correlated group" problem. It suits
    multivariate prediction in landscape/real-estate value modelling.

====================================================================================

#58  Robust Regression
    file: 58_robust_outlier_area_value.xlsx
  >> SCENARIO (narration):
    To prevent outlier parks from distorting the area-value relationship in OLS, we use
    robust regression (Huber-T), which down-weights outliers.
  >> VARIABLE SELECTION:
    - Target: park_value_TL
    - Predictor: area_m2

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Robust intercept = 52,234   vs   OLS intercept = 59,061   (outliers distorted OLS)
    Park value (TL) ~ area

>> COMMENTARY (narration):
    This data had a few outliers -- e.g. parks of extraordinary value/location. Ordinary regression (OLS) takes
    outliers fully into account and is pulled toward them: the intercept is 59,061 in OLS but 52,234 in the robust
    model; the marked gap means outliers distorted OLS. Robust regression (Huber-T) gives less weight to outlying
    observations to preserve the true area-value relationship. In landscape/real-estate value data, where outliers are
    inevitable, robust regression gives a more reliable slope estimate than OLS.

====================================================================================

#59  Quantile Regression
    file: 59_quantile_park_value.xlsx
  >> SCENARIO (narration):
    We examine area's and green ratio's effect on park value separately at the lower/
    middle/upper quantiles of the distribution; quantile regression is appropriate.
  >> VARIABLE SELECTION:
    - Target: park_value_TL
    - Predictors: area_m2, green_ratio  |  Quantiles: 0.1, 0.5, 0.9

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    q = 0.10 slope = 0.0996   |   q = 0.50 (median)   |   q = 0.90  (varying effect)
    Park value ~ area + green ratio

>> COMMENTARY (narration):
    Ordinary regression models only the mean; but area and green ratio may affect park value differently for low-value
    and high-value parks. Quantile regression models the lower (q=0.10, low value), middle (q=0.50) and upper (q=0.90,
    high value) quantiles separately. A different slope across quantiles answers "does area/green have the same effect
    at every value level" -- for instance green ratio may command a higher premium in the most valuable parks. In
    landscape/real-estate valuation, when the different parts of the distribution matter rather than the average,
    quantile regression reveals what mean-based models miss.

====================================================================================

#60  ROC Curve
    file: 60_roc_camera_confidence.xlsx
  >> SCENARIO (narration):
    We evaluate how well a camera-confidence score discriminates "safe park" perception
    with the ROC curve/AUC.
  >> VARIABLE SELECTION:
    - Target (binary): safe_label
    - Score/predictor: camera_confidence_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    AUC = 0.958 (excellent discriminating power)   Youden optimum threshold: 0.53 -> Sensitivity 0.95
    n = 250 (positive 152, negative 98)   safety classifier

>> COMMENTARY (narration):
    We measured the discriminating power of a classifier predicting whether a park is perceived as "safe" from a
    camera-confidence score with the ROC curve. AUC = 0.96 is "excellent": the camera score separates safe from unsafe
    parks almost perfectly. The Youden index gives the best cut-off (threshold), optimising sensitivity and
    specificity; here sensitivity is 0.95 at a 0.53 threshold. ROC/AUC is the standard for reporting a diagnostic/
    classification model's discriminating power; in landscape it is used for safety-perception prediction, smart
    camera/monitoring systems and safe-place design (CPTED) studies.

====================================================================================

#61  TSS
    file: 61_tss_type_distribution.xlsx
  >> SCENARIO (narration):
    We evaluate, with the prevalence-insensitive TSS, the skill of a model predicting a
    species'/habitat's presence from environmental variables.
  >> VARIABLE SELECTION:
    - Target (binary): type_present
    - Predictors: green_ratio, humidity_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    TSS = 0.472 (good)   Sensitivity 0.63   Specificity 0.84   n = 200
    Species/habitat presence prediction

>> COMMENTARY (narration):
    TSS (True Skill Statistic) is a common performance measure in species/habitat distribution models: sensitivity +
    specificity - 1. The model predicting a species'/habitat's presence from environmental variables (green ratio,
    humidity) has TSS = 0.47 -- "good": specificity high (0.84, separates absences well), sensitivity moderate (0.63).
    TSS's advantage is insensitivity to prevalence, so it is preferred in ecology as an alternative to AUC. In
    landscape ecology, SDM and TSS are used for species/habitat suitability mapping and green-infrastructure planning
    -- which areas are suitable for particular species/habitats.

====================================================================================

#62  Confusion Matrix
    file: 62_complexity_park_class.xlsx
  >> SCENARIO (narration):
    We evaluate a park-class classifier's actual vs predicted labels with a confusion
    matrix (accuracy, F1).
  >> VARIABLE SELECTION:
    - Actual: actual_class
    - Predicted: prediction_class

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.94   F1 = 0.923
    Park-class prediction -- actual vs predicted label

>> COMMENTARY (narration):
    We broke down a park-class classifier's predictions with a confusion matrix. With accuracy 94% and F1 = 0.92 the
    model is very successful -- both the overall correct rate is high and the precision/recall balance (F1) is good.
    The confusion matrix reveals what the "accuracy" number hides -- where each park class is misclassified, which
    classes are confused. The F1 score is more informative than accuracy especially with imbalanced classes. In
    landscape it is the indispensable tool for evaluating the real performance of automated park classification/
    typology models (satellite/GIS-based).

====================================================================================

#63  Random Forest
    file: 63_rf_park_type.xlsx
  >> SCENARIO (narration):
    We classify park type from five physical features with Random Forest and inspect
    variable importance.
  >> VARIABLE SELECTION:
    - Target: park_type
    - Features: area_m2, green_ratio, coast_proximity, tree_density, age_year

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.914   (park-type classification)
    Predicted from 5 features (area, green ratio, coast proximity, tree density, age)

>> COMMENTARY (narration):
    We predicted parks' type from five features (area, green ratio, coast proximity, tree density, age) with Random
    Forest -- the vote of hundreds of decision trees. Accuracy is 91%: the features largely classify park type
    correctly. One of Random Forest's most valuable outputs is the "variable importance ranking", telling which
    feature matters most for the type distinction. It is a powerful, easy-to-build machine-learning method for
    nonlinear, interacting relationships; in landscape it is valuable for automatically classifying large-scale park
    inventories (with GIS/remote-sensing support).

====================================================================================

#64  SVM
    file: 64_svm_park_safe.xlsx
  >> SCENARIO (narration):
    We find the best maximum-margin boundary separating safe/unsafe parks with SVM.
  >> VARIABLE SELECTION:
    - Target (binary): safe
    - Features: brightness_lux, camera_coverage, pedestrian_density, crime_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (park safe/not)
    Predicted from 4 variables (brightness, camera coverage, pedestrian density, crime ratio)

>> COMMENTARY (narration):
    We found the best boundary separating safe from unsafe parks with SVM. SVM seeks the maximum-margin separating
    surface between two classes and, via the kernel trick, can model nonlinear boundaries too. Accuracy is 100% --
    brightness, camera coverage, pedestrian density and crime ratio separate safe/unsafe parks perfectly (such high
    accuracy should be kept in mind regarding overfitting). SVM is especially powerful in small-to-medium, high-
    dimensional classification problems. In safe-place design (CPTED) analysis -- safety from lighting/surveillance/
    vitality variables -- it is a strong alternative to Random Forest.

====================================================================================

#65  Gradient Boosting
    file: 65_gb_renewal_urgency.xlsx
  >> SCENARIO (narration):
    We predict parks' renewal urgency from management variables with Gradient Boosting to
    prioritise in asset management.
  >> VARIABLE SELECTION:
    - Target: renewal_urgency
    - Features: age_year, maintenance_score, annual_complaint, annual_budget

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 0.911   (renewal-urgency classification)
    age + maintenance score + annual complaints + annual budget

>> COMMENTARY (narration):
    We predicted parks' renewal urgency (low/medium/high) from age, maintenance score, complaints and budget with
    Gradient Boosting. Boosting adds trees sequentially, each new tree correcting the previous ones' errors; accuracy
    is high at 91.1%. This is critical for asset management: high-renewal-urgency parks can be identified in advance
    and budget/planning prioritised. While Random Forest builds trees in parallel, Boosting builds them sequentially
    (chasing the error); it usually gives higher accuracy but needs careful tuning. It is a valuable decision-support
    tool in landscape asset management and renewal planning.

====================================================================================

#66  K-Means Clustering
    file: 66_kmeans_park_4kume.xlsx
  >> SCENARIO (narration):
    We split parks into natural groups (a typology) by physical features with K-Means; 4 clusters.
  >> VARIABLE SELECTION:
    - Variables: green_ratio, area_m2, tree_density, path_length_m  |  Clusters: 4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.543 (strong cluster structure)
    Park typology from 4 variables (green ratio, area, tree density, path length)

>> COMMENTARY (narration):
    We split parks into four natural groups (a typology) by their physical features with K-Means. Silhouette = 0.54 is
    strong -- the clusters separate clearly, a real structure exists. These four groups likely represent distinct park
    types: e.g. large-green city parks, small-dense neighbourhood parks, tree-rich forest-like parks, and orderly-
    structured parks. K-Means finds natural groups in unlabeled data; building a park typology, grouping similar parks
    and setting region-based management strategy is extremely practical in landscape. The silhouette score confirms
    the number of clusters was chosen well.

====================================================================================

#67  Hierarchical Clustering
    file: 67_hierarchic_5tur_10feat.xlsx
  >> SCENARIO (narration):
    We group units by 10 features with hierarchical clustering (dendrogram); 5 clusters.
  >> VARIABLE SELECTION:
    - Variables: feat_01..feat_10  |  Clusters: 5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 5   Silhouette = 0.145 (weak/overlapping structure)
    5 types/groups from 10 features

>> COMMENTARY (narration):
    We grouped units by 10 features with hierarchical clustering -- the result is a dendrogram (tree). Splitting into
    five clusters gives a silhouette of only 0.15: the clusters overlap, there is no sharp separation. This is common
    in natural/environmental data -- features often change along a continuous gradient rather than in sharp categories.
    A low silhouette means "natural grouping in the data is weak", which is informative (perhaps fewer than 5 groups
    exist). Hierarchical clustering is more explanatory than K-Means for discovering the nested structure of groups and
    the right number of clusters; the dendrogram visualises the similarity hierarchy.

====================================================================================

#68  DBSCAN
    file: 68_dbscan_dog_park.xlsx
  >> SCENARIO (narration):
    We cluster facility (dog-park) locations by density with DBSCAN, separating isolated
    points (noise).
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   Silhouette = 0.871   (location coordinates)
    Dog-park/area locations clustered by density

>> COMMENTARY (narration):
    We clustered dog-park/area locations by geographic density with DBSCAN. Unlike K-Means: we do not give the number
    of clusters in advance, and it flags low-density isolated locations as "noise" (outliers). Four dense groups were
    found (silhouette 0.87, very strong). This shows these facilities cluster in particular regions. DBSCAN is more
    appropriate than K-Means for spatial data with irregularly shaped clusters and outliers (facility locations,
    service points). In landscape it is used in service-accessibility analysis (which regions are dense in dog parks/
    playgrounds, which are empty) and in planning new facility locations.

====================================================================================

#69  PCA
    file: 69_pca_park_6ozellik.xlsx
  >> SCENARIO (narration):
    We reduce six park features to a few summary axes and reveal hidden dimensions with PCA.
  >> VARIABLE SELECTION:
    - Variables: area_m2, green_ratio, tree_count, path_length_m, water_area_pct, satisfaction

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    PC1 = 87.5%   (the first component alone explains most of the variance)
    6 park features (area, green ratio, tree count, path length, water area, satisfaction)

>> COMMENTARY (narration):
    We reduced six park features to a few summary axes with PCA. The first component alone explains 87.5% of the
    variance -- strikingly high. The meaning: these features are very highly correlated, all measuring a single "park
    size/richness" dimension (large parks generally have many trees, long paths, higher satisfaction). PCA is the basic
    tool for summarising multivariate park data, reducing collinearity and revealing hidden dimensions (like general
    park quality). One component being this dominant suggests a composite park-quality index could be built.

====================================================================================

#70  t-SNE
    file: 70_tsne_leaf_5tur.xlsx
  >> SCENARIO (narration):
    We visually explore five plant species defined by 12 leaf measurements on a 2D map with t-SNE.
  >> VARIABLE SELECTION:
    - Variables: measurement_01..measurement_12

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KL Divergence (fit quality) = 0.534 -> good
    12 leaf measurements (5 plant species) reduced to 2D

>> COMMENTARY (narration):
    t-SNE is a nonlinear method that projects high-dimensional leaf-morphology data (12 measurements) into a two-
    dimensional plot. Unlike PCA, it focuses on preserving local neighbourhoods -- similar leaves/species are placed
    close on the map -- ideal for seeing hidden species structure. The KL divergence is 0.53, so the reduction quality
    is good; the five plant species form visible clusters on the map. Important caveat: t-SNE axes and between-cluster
    distances are not interpretable, only the clustering pattern is. In landscape plant-material identification and
    morphological species discrimination it is a powerful visual-exploration tool.

====================================================================================

#71  MDS
    file: 71_mds_park_distance.xlsx
  >> SCENARIO (narration):
    We map the similarity among parks onto two dimensions with multidimensional scaling (MDS).
  >> VARIABLE SELECTION:
    - Variables: green_ratio, area_m2, tree_density, path_length_m, water_area_pct

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Stress (Kruskal-1) = 0.033 (good fit)
    Park similarity map (green ratio, area, tree density, path length, water area)

>> COMMENTARY (narration):
    MDS maps the similarity among parks into two dimensions -- parks with similar features placed near, dissimilar ones
    far. Stress = 0.033 is a "good" fit; this means the park-similarity structure fits two dimensions successfully (low
    stress = reliable map). Parks close on the map have similar physical profiles. MDS is the classic way to visualise
    multivariate similarity/distance data; the stress value honestly tells how much we can trust the map. It is used to
    see park-similarity grouping and typological positioning.

====================================================================================

#72  UMAP
    file: 72_umap_spectral_5tip.xlsx
  >> SCENARIO (narration):
    We explore land-cover types defined by 25 spectral channels on a 2D map with UMAP.
  >> VARIABLE SELECTION:
    - Variables: channel_01..channel_25

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    UMAP dimension reduction (25 spectral channels -> 2D, 5 types)
    n_neighbors and min_dist balance local/global structure

>> COMMENTARY (narration):
    UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure
    better and is faster. We projected area profiles of 25 spectral channels (satellite/hyperspectral) into a two-
    dimensional map; similar spectral signatures (land-cover types) cluster. The n_neighbors parameter tunes the
    local-global balance, and min_dist the tightness of clusters. UMAP is increasingly preferred for discovering hidden
    land-cover classes in high-dimensional remote-sensing/spectral data. In landscape/GIS it is valuable for land-cover
    classification and visualisation.

====================================================================================

#73  Cronbach's Alpha
    file: 73_cronbach_park_quality_21.xlsx
  >> SCENARIO (narration):
    We test the internal consistency (reliability) of a 21-item park-quality scale with
    Cronbach's alpha.
  >> VARIABLE SELECTION:
    - Items: item_01..item_21

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.967 (excellent internal consistency)   21 items
    Park-quality assessment scale

>> COMMENTARY (narration):
    We tested the internal consistency of a 21-item scale measuring park quality with Cronbach's alpha. Alpha = 0.97
    is excellent: the 21 items consistently measure the same underlying concept (park quality), and users answer
    coherently across items. Showing a scale's reliability before analysing it is essential -- low alpha signals the
    items do not measure the same thing and the total score may be meaningless. Above 0.70 is acceptable, above 0.90
    excellent. It is the first quality step of any study developing landscape quality/satisfaction scales. (A very high
    alpha can also hint at item redundancy.)

====================================================================================

#74  Likert Scale Analysis
    file: 74_likert_park_3boyut.xlsx
  >> SCENARIO (narration):
    We examine the reliability and item statistics of a three-dimensional (aesthetic/
    function/safety) 15-item Likert scale.
  >> VARIABLE SELECTION:
    - Items: aesthetic_1..5, function_1..5, safety_1..5

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Cronbach's alpha = 0.760 (acceptable)   15 items (3 dimensions: aesthetic/function/safety)

>> COMMENTARY (narration):
    We examined the overall reliability and item statistics of a 15-item, three-dimensional (aesthetic, function,
    safety) Likert park-evaluation scale. Cronbach's alpha = 0.76 is "acceptable" -- the scale is generally consistent,
    but because three different dimensions are combined, a single alpha can come out a bit low (sub-dimensions should
    be analysed separately). Likert analysis gives item means, distributions and item-total correlations to show
    whether the scale is sound. In landscape it is used for the basic quality control of park-perception/evaluation
    scales.

====================================================================================

#75  Exploratory Factor Analysis (EFA)
    file: 75_efa_18_item_3faktor.xlsx
  >> SCENARIO (narration):
    We explore the hidden factor structure (expected 3 factors) underlying 18 items with EFA.
  >> VARIABLE SELECTION:
    - Items: q01..q18

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    KMO (overall) = 0.900 (Excellent)   -> factor analysis very appropriate
    18 items, expected 3 factors

>> COMMENTARY (narration):
    We explored how many hidden dimensions (factors) underlie 18 scale items with Exploratory Factor Analysis. KMO =
    0.90 is "excellent": the data are very suitable for factor analysis, with high shared variance among items. The
    items group into three factors as expected. EFA reduces many items to a few meaningful dimensions, simplifying the
    scale and revealing the underlying structures (aesthetic/function/environment). It is a core tool of developing
    landscape quality/perception scales and of psychometric validation; KMO and Bartlett's test are reported as the
    pre-check of suitability.

====================================================================================

#76  Intraclass Correlation (ICC)
    file: 76_icc_3uzman_50park.xlsx
  >> SCENARIO (narration):
    We measure the consistency (inter-rater reliability) of three experts evaluating the
    same 50 parks with ICC.
  >> VARIABLE SELECTION:
    - Raters: expert_1, expert_2, expert_3

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ICC(1,1) = 0.880   (Excellent)   3 experts / 50 parks   95% CI ~ (0.82 , 0.93)

>> COMMENTARY (narration):
    We assessed how consistent three experts are in evaluating the same 50 parks with ICC. ICC = 0.88 is "excellent":
    inter-expert agreement is very high -- whichever expert evaluates, the result is similar. This is critical in
    landscape quality assessment: if scoring varied from expert to expert, park rankings and planning decisions would
    be unreliable. ICC is the standard for measuring inter-rater reliability on continuous measurements (scores,
    ratings) (Cohen's Kappa is for categorical data, ICC for continuous). It is the standard way to prove the
    reliability of expert-based landscape evaluations.

====================================================================================

#77  Confirmatory Factor Analysis (CFA)
    file: 77_cfa_12_item_3faktor.xlsx
  >> SCENARIO (narration):
    We test whether the pre-specified three-factor (aesthetic/access/comfort) structure
    fits the data with CFA.
  >> VARIABLE SELECTION:
    - Factors: aesthetic_1..4 ; access_1..4 ; comfort_1..4

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    CFI = 1.00   RMSEA = 0.000   (excellent fit)
    3 factors: aesthetic + access + comfort

>> COMMENTARY (narration):
    While EFA explores hidden structure, CFA tests a pre-specified theory: here we tested whether three distinct
    park-quality dimensions -- "aesthetic", "access" and "comfort" -- fit the data. The fit is excellent (CFI = 1.00,
    RMSEA = 0.000): the items really load onto these three dimensions, the hypothesised scale structure is fully
    confirmed. CFA is part of structural equation modelling (SEM) and is the strongest way to prove scale validity. In
    landscape it is the standard method for demonstrating the validity of multidimensional constructs like park quality
    and place attachment; in evidence-based design research it proves the soundness of scales.

====================================================================================

#78  Survey Means (Complex Sample)
    file: 78_survey_means_income.xlsx
  >> SCENARIO (narration):
    From a stratified, weighted user survey we estimate the weighted mean of monthly
    income and its correct standard error.
  >> VARIABLE SELECTION:
    - Variable: monthly_income_TL  |  Weight: weight  |  Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Monthly income weighted mean = 13,282.61 TL   SE = 185.85   95% CI (12,917 , 13,648)   CV 1.40%

>> COMMENTARY (narration):
    A user survey's participants were selected not at random but with a stratified, weighted design. A simple mean
    would be wrong for such data; complex-sample analysis gives the correct mean income and standard error (Taylor
    linearization) by accounting for the design weights and stratification. Park users' monthly income weighted mean
    is 13,283 TL, with a very narrow confidence interval (CV 1.40%). Profiling park users' socio-economic make-up
    matters in equitable green-space access analysis (who reaches parks). Survey methods are essential for population-
    generalisable, unbiased estimation.

====================================================================================

#79  Survey Frequency (Complex Sample)
    file: 79_survey_freq_use.xlsx
  >> SCENARIO (narration):
    We estimate the proportions of park-use-frequency categories accounting for design
    weights and stratification.
  >> VARIABLE SELECTION:
    - Variable: use_frequency  |  Weight: weight  |  Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    "every day" proportion = 0.083 (SE 0.014)   |   "none" proportion = 0.440 (SE 0.025)
    Park-use frequency -- stratified + weighted design

>> COMMENTARY (narration):
    We estimated the proportion of park-use-frequency categories among users, accounting for the complex sample design.
    Daily users are 8%, never-users 44% (with design-based standard errors). This is striking -- nearly half the
    population never uses parks; this may point to an access/attractiveness problem. A simple percentage would give the
    wrong standard error by ignoring stratification and weights; survey frequency analysis produces a confidence
    interval reflecting the true uncertainty. It is the correct way to estimate use proportions (park use, green-space
    access) in a population in a design-consistent manner -- the basis of green-space policy monitoring.

====================================================================================

#80  Survey Total (Complex Sample)
    file: 80_survey_total_green_prediction.xlsx
  >> SCENARIO (narration):
    We estimate the total green area (and its uncertainty) by scaling from the sample to
    the whole city; complex-sample total is appropriate.
  >> VARIABLE SELECTION:
    - Variable: green_area_m2  |  Weight: weight  |  Stratum: stratum

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Total green area = 6,511,244 m2   SE = 91,997   95% CI (6,330,141 , 6,692,347)
    Stratified design (Taylor SE)

>> COMMENTARY (narration):
    We estimated the total green area in a city/region using the sampling design and weights: about 6.5 million m2 (95%
    CI 6.33 - 6.69 million). Scaling from the sampled areas up to the whole city requires correct weighting -- each
    sampled unit carries weight proportional to the population fraction it represents. Survey total analysis does
    exactly this and expresses uncertainty with the standard error. It is the official-statistics method used to
    estimate landscape planning indicators such as total green area, per-capita green area and the urban green-
    infrastructure inventory from a sample.

====================================================================================

#81  Survey Regression (Complex Sample)
    file: 81_survey_reg_income.xlsx
  >> SCENARIO (narration):
    From a stratified/weighted survey we predict monthly income from age and education
    with design-based standard errors.
  >> VARIABLE SELECTION:
    - Target: monthly_income_TL  |  Predictors: age, education_year  |  Weight: weight  |  Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.731   age coefficient = 77.50   p < .001 ***
    Monthly income ~ age + education years   (stratified + weighted)

>> COMMENTARY (narration):
    We modelled park users' monthly income from age and education years, accounting for the complex sample design
    (strata, weights). The age effect is significant (coefficient 77.50, p < .001) and the model explains 73% of the
    variance. Design-based standard errors are more accurate than the (often too optimistic) errors of simple OLS.
    This is the valid way to relate the park-user profile to socio-economic variables -- important in equity analysis
    of green-space access. When estimating relationships from national/urban survey data, survey regression is
    necessary for valid inference; otherwise p-values would mislead.

====================================================================================

#82  Survey Logistic Regression
    file: 82_survey_logistic_repeat.xlsx
  >> SCENARIO (narration):
    We predict the probability of repeat visit (binary) from satisfaction/age with the
    complex sample design.
  >> VARIABLE SELECTION:
    - Target (binary): repeat_visit  |  Predictors: age, satisfaction  |  Weight: weight  |  Stratum: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    satisfaction: OR = 1.966   p < .001 ***
    Repeat visit ~ age + satisfaction   (stratified + weighted)

>> COMMENTARY (narration):
    Predicting a user's probability of revisiting the park from satisfaction, we handled both the binary outcome
    (logistic) and the complex sample design together. The odds ratio for satisfaction is 1.966 (p < .001):
    satisfaction roughly doubling the odds of revisiting -- satisfaction's strong role in park loyalty. Because the
    design weights and stratification are accounted for, the OR and confidence interval are design-consistent and
    valid. Survey logistic regression is the correct way to model binary outcomes (repeat visit, loyalty) from complex
    user surveys; it is valuable for identifying the factors that build park loyalty.

====================================================================================

#83  Generalized Additive Model (GAM)
    file: 83_gam_PM25_temperature.xlsx
  >> SCENARIO (narration):
    We model the nonlinear relationship between air pollution (PM2.5) and temperature with
    smooth curves (splines); GAM is appropriate.
  >> VARIABLE SELECTION:
    - Target: PM25_ugm3  |  Predictor: s(temperature_C)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (explained) = 0.790
    PM2.5 (air pollution) ~ s(temperature)  (smooth function)

>> COMMENTARY (narration):
    GAM is a flexible generalisation of linear regression: it models a variable's effect with "smooth curves"
    (splines) rather than a straight line. The relationship between PM2.5 (fine-particulate pollution) and temperature
    is nonlinear -- pollution can peak at certain temperatures (inversion, seasonal effects). GAM learns this curved
    pattern from the data without assuming its shape in advance; the model explains 79% of the variance. It preserves
    interpretability while adding flexibility: we can read from the curve at which temperature pollution rises. In
    landscape, for modelling green infrastructure's air-quality/climate-regulation ecosystem service -- nonlinear
    environmental relationships -- it is an ideal method.

====================================================================================

#84  Discriminant Analysis
    file: 84_diskriminant_3park_5feat.xlsx
  >> SCENARIO (narration):
    We find the linear combinations best separating three park types and classify new
    parks; discriminant analysis is appropriate.
  >> VARIABLE SELECTION:
    - Target: park_type  |  Features: area_m2, green_ratio, tree_density, coast_proximity, path_length_m

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Accuracy = 1.000   (3 park types)
    Separation from 5 features (area, green ratio, tree density, coast proximity, path length)

>> COMMENTARY (narration):
    Discriminant analysis finds the linear combinations of variables that best separate the groups (three park types)
    -- a bit like a classification-focused PCA. Accuracy is 100%: five features separate the three park types
    perfectly. Discriminant analysis both classifies and answers "which feature matters most for the separation"
    (discriminant function loadings). It resembles logistic regression but is the classic choice for multiple groups
    when variables are normally distributed. In landscape it is used in park-type classification and assigning new
    parks to an existing typology; it provides a basis for planning standards and type-based management.

====================================================================================

#85  Conditional Logit
    file: 85_conditional_logit_transport.xlsx
  >> SCENARIO (narration):
    We model which transport mode users take to the park (discrete choice) from duration
    and cost; conditional logit is appropriate.
  >> VARIABLE SELECTION:
    - Chooser: participant_id  |  Chosen: chosen  |  Predictors: duration_min, cost_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Pseudo R2 (McFadden) = 0.803  (very high)
    Transport-mode choice ~ duration (min) + cost (TL)

>> COMMENTARY (narration):
    We examined which transport mode (walking, bicycle, public transport, car) users take to the park with a
    conditional logit model -- each user selects one from a choice set. McFadden pseudo R2 = 0.80 is very high
    (0.2-0.4 is already strong for choice models): duration and cost explain the mode choice almost completely. The
    model shows users prefer shorter, lower-cost modes. This method is the basis of transport and accessibility
    analysis; in landscape it is powerful for modelling access modes to parks and policies promoting green/active
    travel (walking, cycling).

====================================================================================

#86  Kaplan-Meier Survival
    file: 86_km_street_agaci.xlsx
  >> SCENARIO (narration):
    We examine street trees' survival time by type and compare with the log-rank test;
    for censored "time-to-event" data Kaplan-Meier is appropriate.
  >> VARIABLE SELECTION:
    - Duration: duration_year  |  Event: event  |  Grouping: type

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Median survival = 8.41 years   (grouped by tree type)
    Log-rank test compares groups

>> COMMENTARY (narration):
    We analysed street trees' survival time -- by type -- with Kaplan-Meier. Median survival is 8.41 years (reflecting
    the harshness of urban street conditions -- heat, pollution, limited root space). Kaplan-Meier analyses "time-to-
    event" data -- correctly using not-yet-dead (censored) trees too -- and tests type differences in survival with the
    log-rank test. In landscape, street-tree survival is the basic indicator of urban-afforestation success and type-
    condition suitability. Determining which type survives longer in urban conditions is critical for species
    selection.

====================================================================================

#87  Cox Proportional Hazards
    file: 87_cox_death_hazard.xlsx
  >> SCENARIO (narration):
    We relate street-tree death risk to environmental variables and obtain hazard ratios
    (HR); Cox regression is appropriate.
  >> VARIABLE SELECTION:
    - Duration: duration_year  |  Event: event  |  Covariates: age_year, stress_score, moisture_pct

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 147 (58.8%)   Concordance = 0.570
    stress_score: HR = 1.009 p = 0.106 (not significant)   (with age, moisture covariates)

>> COMMENTARY (narration):
    We modelled street trees' death risk from environmental/physiological variables (age, stress score, moisture) with
    Cox regression. In this model the stress score is borderline non-significant (HR = 1.009, p = 0.106), concordance
    0.57 indicating modest discrimination. This is informative too: tree death is partly explained by these variables,
    with other factors (root damage, infrastructure conflict, vandalism) also at play. Cox regression is the gold-
    standard survival model relating "time-to-event" to several predictors; it gives hazard ratios (HR). Relating
    urban tree death risk to environmental conditions is valuable for resistant species selection and planting-site
    improvement.

  >> ADVANCED PARAMETERS (optional in the form — what they do):
    - Proportional-hazards assumption test (Schoenfeld): per-covariate chi-square/p; significant = PH violated, consider a time-varying effect.

====================================================================================

#88  Parametric Survival (Weibull AFT)
    file: 88_aft_weibull_omur.xlsx
  >> SCENARIO (narration):
    We examine tree lifespan with a parametric AFT model that fits the survival curve
    with a Weibull distribution.
  >> VARIABLE SELECTION:
    - Duration: duration_year  |  Event: event  |  Covariate: maintenance_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Distribution = Weibull   Median survival = 12.22 years
    Survival ~ maintenance score

>> COMMENTARY (narration):
    We analysed street trees' lifespan with a parametric Weibull model. While Cox is semi-parametric, Weibull models
    the whole survival curve with a specific mathematical distribution -- if the data fit it, it gives more efficient
    estimates and allows extrapolation. Median lifespan is 12.22 years. The AFT (accelerated failure time)
    interpretation is intuitive: how many times a factor like maintenance lengthens/shortens a tree's lifespan. When
    the time distribution is known, parametric models give more powerful and predictable results than Cox. Urban tree-
    lifespan estimation is valuable in tree asset management and renewal planning (how often to replace trees).

====================================================================================

#89  Competing Risks
    file: 89_competing_risks_3neden.xlsx
  >> SCENARIO (narration):
    We examine a tree's mutually exclusive causes of death (disease/stress/damage) in a
    competing-risks framework with a separate cumulative incidence per cause.
  >> VARIABLE SELECTION:
    - Duration: duration_year  |  Event type: death_cause (0=censored)  |  Covariate: age_planting

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Censored (0) = 59   |   3 different causes of death (death_cause types)
    Separate cumulative incidence (CIF) and risk for each event type

>> COMMENTARY (narration):
    A street tree can die from more than one cause: disease, drought/stress, or mechanical damage (construction,
    accident) -- and these are mutually exclusive (once one happens the others cannot). Classic survival handles this
    incorrectly; competing-risks analysis produces a separate cumulative incidence for each cause of death. In the data
    59 trees are still alive (censored), the rest died of different causes. This framework correctly answers questions
    like "which cause of death does drought/age increase". In urban-tree survival research, when there are different
    causes of death, competing-risks analysis is essential -- otherwise one cause's risk is overstated and the wrong
    protective intervention is chosen.

====================================================================================

#90  Time-Dependent Cox
    file: 90_tvcox_tree_stress.xlsx
  >> SCENARIO (narration):
    Using time-varying stress score (start-stop), we model the effect of current status
    on death risk; time-dependent Cox is appropriate.
  >> VARIABLE SELECTION:
    - ID: tree_id  |  Start/Stop: start, end  |  Event: event  |  Covariate: stress_score

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    stress_score: HR = 1.002   95% CI (0.978 , 1.028)   p = 0.851 (non-significant)
    Time-varying stress-score covariate

>> COMMENTARY (narration):
    A tree's stress score changes over time (drought periods, season); standard Cox assumes it constant, time-dependent
    Cox uses the current stress score in each time interval (start-stop format). In this model the effect of current
    stress on death risk is non-significant (HR = 1.002, p = 0.851) -- in this data the momentary stress change does
    not predict death timing. A non-significant result is information too. Time-dependent Cox is the correct method
    when "the current, continuously monitored condition matters, not just the baseline"; it is the standard for
    modelling time-varying stress/health measurements in longitudinal tree monitoring.

====================================================================================

#91  Survey Cox (Survey PHREG)
    file: 91_survey_phreg_project.xlsx
  >> SCENARIO (narration):
    We model project completion with survey-PHREG, accounting for the weighted/clustered
    sample design.
  >> VARIABLE SELECTION:
    - Duration: duration_month  |  Event: completed  |  Covariate: budget_TL  |  Weight: weight  |  Cluster: cluster_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.511   Weighted/clustered survival (budget covariate, region cluster)
    Project completion ~ budget

>> COMMENTARY (narration):
    We combined survival analysis with a complex sample design: landscape projects come from a population sampled with
    weights and clustered by regions; the "event" is project completion. Survey-PHREG runs the Cox model with design
    weights and cluster-robust standard errors, so the hazard ratios and confidence intervals generalise to the
    population. Concordance 0.51 means budget alone is a weak discriminator in this model. In landscape project
    management, when there is "time-to-completion" data, ignoring the design biases the estimates; survey-PHREG
    provides valid inference.

====================================================================================

#92  Interval-Censored Survival
    file: 92_interval_censored_disease.xlsx
  >> SCENARIO (narration):
    The event (disease onset) time lies within an interval between two checks; we handle
    this with the Turnbull algorithm.
  >> VARIABLE SELECTION:
    - Lower bound: left_censor  |  Upper bound: survival_censor

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Events = 28   Median survival = 8.0 years   (Turnbull algorithm)
    Disease-onset time known within [lower, upper] interval

>> COMMENTARY (narration):
    In urban-tree monitoring we rarely know the exact time of an event (disease onset): we check the tree periodically,
    find it healthy at one check and diseased at the next -- the disease started somewhere between the two dates.
    Interval-censored survival (Turnbull algorithm) handles exactly this uncertainty, avoiding fixing the event to an
    arbitrary date. Median time to disease is 8 years. By the nature of periodic inspection (annual check), event times
    are always within an interval; forcing this data to a single point biases the result. The interval-censored method
    matches the reality of periodic tree-health monitoring data.

====================================================================================

#93  Frailty Cox
    file: 93_frailty_cox_neighborhood.xlsx
  >> SCENARIO (narration):
    For neighborhood-clustered tree survival, we use frailty Cox, which adds a shared
    neighborhood-level random effect ("frailty").
  >> VARIABLE SELECTION:
    - Duration: duration_year  |  Event: event  |  Covariate: age_planting  |  Cluster (frailty): neighborhood_id

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Concordance = 0.554   Neighborhood-clustered shared frailty (random effect)
    Death ~ planting age  +  (neighborhood frailty)

>> COMMENTARY (narration):
    Street trees are clustered within neighborhoods; trees in the same neighborhood share unmeasured common factors
    (microclimate, maintenance regime, pollution, infrastructure) and carry similar death risk. Frailty Cox adds a
    shared "frailty" (random effect) for each neighborhood to model this clustering -- the mixed-model version of
    survival analysis. Concordance 0.55. Accounting for neighborhood-level hidden differences gives both correct
    standard errors and information on "how much heterogeneity exists in tree survival among neighborhoods" -- showing
    which neighborhoods are tree-friendly/hostile. It is the correct method for clustered urban-tree survival data
    (neighborhoods, streets).

====================================================================================

#94  Time Series Analysis
    file: 94_ts_monthly_visitor.xlsx
  >> SCENARIO (narration):
    We reveal the structure -- trend, seasonality, stationarity -- of the monthly park-
    visitor series.
  >> VARIABLE SELECTION:
    - Date: date  |  Value: visitor_monthly

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    60 observations (monthly)   ADF p = 0.755 -> NOT stationary   Trend = increasing   Seasonality detected

>> COMMENTARY (narration):
    We examined a monthly park-visitor series. The ADF test does not find it stationary (p = 0.755) -- the mean changes
    over time, there is a trend (increasing -- perhaps rising awareness/use) and strong seasonality is present (summer
    peak, winter low). This pre-diagnosis is critical: most time-series models (like ARIMA) require stationarity, so
    differencing/transformation may be needed first. Time-series analysis reveals the structure of long monitoring data
    -- trend, seasonality, stationarity -- and tells which modelling steps are required. It is the foundation for
    analysing how park-use data (visitors, occupancy) evolve over time.

====================================================================================

#95  STL Decomposition
    file: 95_stl_PM25_daily.xlsx
  >> SCENARIO (narration):
    We decompose the daily PM2.5 series into trend + seasonal + residual components; STL is appropriate.
  >> VARIABLE SELECTION:
    - Date: date  |  Value: PM25_ugm3

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Period (seasonal length) = 7 (daily -> weekly cycle)
    Series decomposed into trend + seasonal + residual components

>> COMMENTARY (narration):
    STL (Seasonal-Trend decomposition using Loess) splits a daily PM2.5 (air pollution) series into three components:
    long-term trend, a recurring cycle (7-day = weekly) and the remaining residual. The weekly period is meaningful --
    air pollution fluctuates regularly across the week (traffic/activity pattern). This decomposition clarifies "is
    pollution generally rising, or just fluctuating weekly". STL is the most intuitive way to interpret seasonal/
    cyclical environmental series (pollution, temperature), separating trend from cycle to see green spaces' long-term
    effect on air quality.

====================================================================================

#96  ARIMA Forecast
    file: 96_arima_monthly_budget.xlsx
  >> SCENARIO (narration):
    We model the monthly park budget with ARIMA and forecast the future.
  >> VARIABLE SELECTION:
    - Date: date  |  Value: monthly_budget_TL

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    ARIMA model fitted   AIC = 2438.85
    Monthly park budget (TL) -- future forecast

>> COMMENTARY (narration):
    We fitted an ARIMA model to a monthly park maintenance/management budget series and forecast the future. ARIMA
    combines the series' own past values (AR), past errors (MA) and differencing (I, to make it stationary). With AIC =
    2439 the best model was selected. Forecasts rest on the observed trend and seasonal autocorrelation structure and
    come with an uncertainty band. Budget forecasting matters for landscape management: anticipating future
    maintenance/renewal spending enables financial planning and resource allocation. ARIMA is the classic forecasting
    method for financial/management series with seasonality and trend.

====================================================================================

#97  Exponential Smoothing (Holt-Winters)
    file: 97_ets_monthly_visitor_yazpik.xlsx
  >> SCENARIO (narration):
    We forecast the strongly seasonal (summer-peaked) visitor series with Holt-Winters
    exponential smoothing.
  >> VARIABLE SELECTION:
    - Date: date  |  Value: visitor

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Method = Holt-Winters (additive seasonal, period 12)   AIC = 831.93
    Monthly visitor count forecast (summer-peaked)

>> COMMENTARY (narration):
    Holt-Winters exponential smoothing estimates the level + trend + seasonality components in a weighted way (giving
    more weight to the recent past). It captured the strong seasonal pattern of the monthly visitor series (summer
    peak) (AIC = 832). Park visitorship typically shows strong seasonality (high in summer, low in winter); Holt-
    Winters is often very successful for such series and is more intuitive than ARIMA. In landscape/recreation demand
    forecasting (visitors, occupancy) it is an easy-to-build, interpretable alternative; it is valuable for seasonal
    staffing/maintenance planning.

====================================================================================

#98  Mann-Kendall Trend + Sen's Slope
    file: 98_mann_kendall_temperature_trend.xlsx
  >> SCENARIO (narration):
    We test the long-term temperature trend (urban heat island/climate) with the
    nonparametric Mann-Kendall + Sen's slope.
  >> VARIABLE SELECTION:
    - Value: mean_temperature_C

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    p < .001 ***   Trend = INCREASING
    Mean temperature long-term trend

>> COMMENTARY (narration):
    We tested whether mean temperature shows a significant long-term trend with the Mann-Kendall test -- a
    nonparametric, outlier- and distribution-robust trend test. The result is highly significant (p < .001) and the
    direction is increasing: temperature is rising consistently over the years. This is a concrete indicator of the
    urban heat island and climate change. It is critical for landscape architecture: rising temperature makes green
    infrastructure's cooling service (shade, evapotranspiration) even more important -- the rationale for climate-
    adaptive design. Mann-Kendall + Sen's slope is the gold standard for long-term trend detection in climate/
    environmental series.

====================================================================================

#99  Anomaly Detection
    file: 99_anomali_park_inventory.xlsx
  >> SCENARIO (narration):
    We detect extraordinary records (multivariate outliers) in the park inventory.
  >> VARIABLE SELECTION:
    - Variables: area_m2, green_ratio, visitor_weekly, satisfaction

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of anomalies = 24   (IQR-based multivariate outlier detection)
    Extraordinary records in the park inventory

>> COMMENTARY (narration):
    We found the extraordinary records in the park inventory (area, green ratio, visitors, satisfaction) with anomaly
    detection -- 24 parks were flagged as deviating from the normal pattern. These anomalies can point to real
    situations: an extraordinarily large/green park, an unexpectedly low-satisfaction park (a problem), or a data-entry
    error. Anomaly detection automatically picks out the few special cases that need attention from large park
    databases. In landscape management it is extremely practical for early warning (detecting problem parks),
    prioritising intervention and data-quality control.

====================================================================================

#100  Variance Components (VARCOMP)
    file: 100_varcomp_city_neighborhood_park_satisfaction.xlsx
  >> SCENARIO (narration):
    We partition satisfaction variability by which spatial level (city/neighborhood/park)
    it comes from with nested variance components (REML).
  >> VARIABLE SELECTION:
    - Dependent: satisfaction  |  Factors: city (upper), neighborhood (nested in city)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    city = 40.2%   neighborhood (nested in city) = 13.4%   residual = ~46.3%
    Nested: city > neighborhood > park (REML)

>> COMMENTARY (narration):
    We partitioned the variability of park satisfaction by which spatial level -- city, neighborhood or park -- it comes
    from with variance-components analysis: city and neighborhood are nested. The result: 40.2% of the variance comes
    from between-city differences, 13.4% from (within-city) between-neighborhood differences, ~46% from park-level/
    residual variation. So the largest structural source is the city level -- cities differ markedly in satisfaction
    (urban-policy/climate/culture differences). This is a critical inference in multi-level design: it shows where the
    variability concentrates and at which level (city policy, local neighborhood or single park) an intervention should
    be targeted.

====================================================================================

#101  Bayesian t-Test
    file: 101_bayesian_t_test_new_old_design.xlsx
  >> SCENARIO (narration):
    We evaluate the satisfaction difference of new vs old design with the Bayes Factor
    (strength of evidence).
  >> VARIABLE SELECTION:
    - Dependent: satisfaction  |  Group: design_group

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    BF10 = 8.32e+67 (decisive evidence)   Cohen's d = 4.95
    New design vs old design -- satisfaction

>> COMMENTARY (narration):
    We tested whether a new design approach gives different satisfaction from the old with a Bayesian t-test. The Bayes
    Factor BF10 = 8.32e+67 -- astronomically large, "decisive evidence": the data support the hypothesis of a
    difference unimaginably more than no difference. Cohen's d = 4.95 makes the effect extraordinarily large. The new
    design raised satisfaction overwhelmingly. Unlike a classic p-value, the Bayes Factor directly measures the
    strength of evidence for both H1 and H0 and distinguishes "absence of evidence" from "evidence of absence". The new
    design's effect is not just statistical but practically overwhelming -- the Bayesian framework shows this
    convincingly.

====================================================================================

#102  Bayesian Correlation
    file: 102_bayesian_correlation_green_satisfaction_BF10.xlsx
  >> SCENARIO (narration):
    We assess the strength and evidence factor (BF10) of the green-area-ratio - satisfaction
    relationship with Bayesian correlation.
  >> VARIABLE SELECTION:
    - Variable 1: green_area_ratio  |  Variable 2: satisfaction

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    r = 0.823 (very strong)   BF10 = 4.47e+22 (decisive evidence)
    Green-area ratio - satisfaction relationship

>> COMMENTARY (narration):
    We assessed the relationship between green-area ratio and park satisfaction with Bayesian correlation: r = 0.82 is
    very strong and the Bayes Factor BF10 = 4.47e+22 makes the evidence for the relationship's existence decisive. As
    green ratio rises, satisfaction rises markedly -- quantified evidence of landscape architecture's core principle
    (green = satisfaction/well-being). Unlike classic correlation, Bayesian correlation expresses the strength of the
    relationship with a probability distribution and an evidence factor -- also answering "how sure are we". Such a
    large BF10 says the probability that the relationship is coincidental is vanishingly small. It is a powerful
    framework for quantifying green space's effect on user satisfaction.

====================================================================================

#103  Bayesian ANOVA
    file: 103_bayesian_anova_design_region.xlsx
  >> SCENARIO (narration):
    We examine design style's and region's effect on satisfaction with Bayesian ANOVA,
    reporting a Bayes Factor per effect.
  >> VARIABLE SELECTION:
    - Dependent: satisfaction  |  Factor 1: design_style  |  Factor 2: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    design_style: BF10 = 7.66e+06 (decisive evidence)
    Design style x region -> satisfaction

>> COMMENTARY (narration):
    We examined the effect of design style and region on satisfaction with Bayesian ANOVA. The design-style effect is
    overwhelming: Bayes Factor 7.66e+06 -- "decisive evidence". So different design styles differ markedly in
    satisfaction. Bayesian ANOVA gives a separate Bayes Factor for each effect and interaction, answering "which factor
    really matters" more informatively than classic ANOVA -- and can even provide "evidence of absence" for non-
    significant effects. In landscape experiments (design x region) it is a powerful way to evaluate factor effects
    with evidence; it provides the Bayesian basis for evidence-based design decisions.

====================================================================================

#104  Bayesian Hierarchical Model
    file: 104_hierarchical_bayesian_city_random.xlsx
  >> SCENARIO (narration):
    In multi-city data we predict satisfaction from green ratio with a city random effect
    using a Bayesian hierarchical model.
  >> VARIABLE SELECTION:
    - Dependent: satisfaction  |  Group: city  |  Predictor: green_ratio

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Group (city) random variance = 0.0604 (33.5%)   residual = 0.1202 (66.5%)   ICC = 0.335
    Satisfaction ~ green ratio  +  (1 | city)   (Bayesian)

>> COMMENTARY (narration):
    We analysed multi-city data (parks within cities) with a Bayesian hierarchical model -- the Bayesian counterpart of
    LMM. The city-level variance is 33.5% of the total (ICC = 0.335); a significant part of the variability comes from
    between-city differences, most from the park level. So satisfaction depends on both the city and the individual
    park. The Bayesian hierarchical model's strength: it expresses both within- and between-group uncertainty with full
    probability distributions and balances cities with few observations via "partial pooling". In multi-city/multi-
    centre landscape studies it is the modern way to model both local and general effects.

====================================================================================

#105  Spatial Lag Model (SAR)
    file: 105_spatial_sar_park_value_spatial.xlsx
  >> SCENARIO (narration):
    We model park value with SAR, which includes the effect of neighbour values (spatial
    lag, rho).
  >> VARIABLE SELECTION:
    - Dependent: park_value_TL  |  Predictors: area_m2, green_ratio  |  Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    rho (spatial lag) = 0.29   z = 3.19   p = 0.001   Pseudo R2 = 0.81
    Park value ~ area + green ratio  +  neighbour park value

>> COMMENTARY (narration):
    Modelling park value, we accounted for spatial dependence -- the tendency of nearby parks to be similar in value.
    The SAR (spatial autoregressive) model links a park's value to its neighbours' value too (rho). Here rho = 0.29,
    significantly positive (z = 3.19, p = 0.001): the neighbour effect is real and moderate -- a park's value is
    influenced by neighbouring parks/surroundings (neighborhood prestige, access). The model explains 81% of the
    variance. Using a spatial model is critical, because ordinary regression gives spurious significant results when
    spatial autocorrelation is present. SAR is the right way to model urban park/real-estate values with spread and
    neighbour effects.

====================================================================================

#106  Spatial Error Model (SEM)
    file: 106_spatial_error_use_residual.xlsx
  >> SCENARIO (narration):
    We examine park value with the spatial error model, which captures the correlation
    induced by unmeasured spatial factors through the residuals (lambda).
  >> VARIABLE SELECTION:
    - Dependent: park_value_TL  |  Predictors: area_m2, green_ratio  |  Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    lambda (spatial error) = 0.41   z = 3.40   p = 0.001   Pseudo R2 = 0.79
    Park value ~ variables  +  spatial error structure

>> COMMENTARY (narration):
    The spatial error model addresses a different kind of spatial dependence than SAR: not a direct effect of neighbour
    values, but correlation induced through the residuals by unobserved (omitted) spatial factors. Lambda = 0.41 (z =
    3.40, p = 0.001) shows this spatial error structure is significant -- there are unmeasured common spatial factors
    (neighborhood infrastructure, socio-economic context, transport). Which spatial model is appropriate (SAR or SEM)
    depends on whether the effect comes from neighbour values or from unmeasured factors. SEM cleans out these
    unmeasured spatial confounders to estimate the true effect of the variables on park value more accurately.

====================================================================================

#107  Geographically Weighted Regression (GWR)
    file: 107_gwr_local_park_value.xlsx
  >> SCENARIO (narration):
    We map how the park-value relationship varies in space (local coefficients) with GWR.
  >> VARIABLE SELECTION:
    - Dependent: park_value_TL  |  Predictors: area_m2, green_ratio  |  Coordinates: lat, lon

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    R2 = 0.831   (local/location-specific regression)
    Spatial variation of the park-value relationship

>> COMMENTARY (narration):
    Ordinary regression assumes a single park-value relationship for the whole city; but this relationship can vary in
    space -- in one region green ratio may strongly affect value while in another it is weak. GWR (Geographically
    Weighted Regression) captures this spatial heterogeneity by fitting a separate local regression at each location
    (R2 = 0.83). The output is a map of coefficients: a surface showing where green ratio/area's effect on park value
    is strong and where it is weak. This tests "is the relationship the same everywhere" and sets local investment/
    renewal priorities. In urban-landscape value analysis, GWR reveals what the global model hides for mapping spatial
    inequalities and place-specific value.

====================================================================================

#108  Nested Mixed Model (Nested LMM)
    file: 108_nested_lmm_city_park_visit.xlsx
  >> SCENARIO (narration):
    In a nested design (city > park > visit) we test each level's contribution separately;
    nested LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent: satisfaction  |  Replication: visit  |  Upper group: city  |  Lower group: park_no

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    city (P): F = 9.06, p < .001 ***   |   park_no [within city] (F): F = 8.30, p < .001 ***
    visit (R): F = 2.98, p = 0.038 *

>> COMMENTARY (narration):
    In this design the measurements are nested: city > park > visit. The nested mixed model tests each level's
    contribution to the variance separately. The results are significant at all levels: there are real differences
    between cities (F = 9.06), between parks within a city (F = 8.30), and between visits (F = 2.98). Mixing up the
    levels in nested data (e.g. ignoring the city) produces spurious significance or wrong standard errors. Nested LMM
    is the correct framework for hierarchical urban designs (city/park/visit, region/neighborhood/unit), separating
    the genuine contribution of each scale; it is the basis of multi-level urban green-space analysis.

====================================================================================

#109  Crossed Mixed Model (Crossed LMM)
    file: 109_crossed_lmm_design_region.xlsx
  >> SCENARIO (narration):
    In a crossed design (each design in each region) we examine main effects and the
    design x region interaction; crossed LMM is appropriate.
  >> VARIABLE SELECTION:
    - Dependent: satisfaction  |  Factor A: design_style  |  Factor B: region

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    design style: F = 1.23, p = 0.356 (non-significant)   |   region: F = 0.69, p = 0.584 (non-significant)
    design x region interaction: F = 11.62, p < .001 *** (SIGNIFICANT)

>> COMMENTARY (narration):
    Unlike nesting, here two factors are crossed: each design style was applied in each region (fully factorial). A
    striking result: both main effects are non-significant (design: p = 0.356, region: p = 0.584), but the interaction
    is highly significant (F = 11.62, p < .001). This is a classic case of a "pure interaction": design and region
    alone say nothing, but TOGETHER -- specific design-region combinations -- they create a strong effect. So a design
    style can be very successful in one region and fail in another; "there is no single design that suits everywhere".
    Looking only at the main effects and saying "nothing is significant" would be a major error. The crossed mixed
    model is the right way to capture the design×region interaction in context-sensitive landscape design.

====================================================================================

#110  KDE -- Kernel Density (Map)
    file: 110_KDE_park_distribution_density.xlsx
  >> SCENARIO (narration):
    We visualise parks' geographic density with a kernel-density (KDE) heat map and spot
    green-space access gaps.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon  |  Value: park_value_TL  (Map tab)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    95 park points   mean park value = 478,473 TL
    Park-distribution density map via kernel density estimation

>> COMMENTARY (narration):
    On the Map tab, we render the geographic density of parks into a heat map with KDE (Kernel Density Estimation). KDE
    turns points into a smooth density surface: we can see where parks are dense and where there are gaps (green-space
    deficits). The distribution of 95 parks shows how urban green infrastructure concentrates geographically. This is
    used both to detect green-space access gaps (which neighborhoods are park-less) -- critical for the "15-minute
    city" / equitable green-space distribution -- and to see park-dense regions. It is the basic mapping tool for
    visualising spatial park data.

====================================================================================

#111  Hexbin (Map)
    file: 111_Hexbin_visitor_distribution.xlsx
  >> SCENARIO (narration):
    We aggregate dense visitor points into hexagon cells for a readable density map.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon  |  Density: density_interval  (Map tab)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    150 visitor/unit points   aggregated into hexagonal cells
    Each cell shows the point density within it by colour

>> COMMENTARY (narration):
    The hexbin map aggregates many points into hexagon cells and shows each cell's density by colour -- a clear
    solution, alternative to KDE, when thousands of points overlap into an unreadable mess. We rendered 150 visitor
    points into a hexagonal grid to visualise where use density is highest. Hexagonal cells carry less edge-bias than
    squares and represent neighbour relations more evenly. It is the practical way to turn dense spatial use data
    (visitor location, activity point) into a readable density map; it is used to identify heavily used areas within a
    park.

====================================================================================

#112  Moran's I (Map)
    file: 112_Morans_I_park_value_autocorrelation.xlsx
  >> SCENARIO (narration):
    We test the spatial autocorrelation of park value (are nearby parks similar?) with Moran's I.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon  |  Value: park_value_TL  (Map tab)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Moran's I = 0.555   z = 9.94   p < .001
    Spatial autocorrelation of park value (strong positive)

>> COMMENTARY (narration):
    Moran's I is the global spatial-autocorrelation statistic measuring "do nearby parks have similar value". The
    result is strongly positive (I = 0.56, z = 9.94, p < .001): park value is spatially clustered -- high-value parks
    lie near each other, and so do low ones. This is not chance; neighborhood prestige, socio-economic context and
    urban development make neighbouring parks similar. This finding is a concrete indicator of spatial inequality in
    urban landscape and is methodologically important: if spatial autocorrelation exists, ordinary statistics are wrong
    (hence the spatial models in #105-107). Moran's I is the starting diagnostic of spatial analysis.

====================================================================================

#113  Getis-Ord Gi* (Map)
    file: 113_Getis_Ord_high_value_hotspot.xlsx
  >> SCENARIO (narration):
    We map park value's hot/cold spots (where it concentrates) with Getis-Ord Gi*.
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon  |  Value: park_value_TL  (Map tab)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    17 hot spots (high-value cluster)   27 cold spots (low-value cluster)   n = 68
    Getis-Ord Gi* local statistic (z > 1.96 significant)

>> COMMENTARY (narration):
    While Moran's I reports the overall clustering of the whole city, Getis-Ord Gi* shows WHERE the hot/cold spots are,
    location by location. 17 parks are statistically significant "hot spots" (high value surrounded by high value --
    prestigious clusters), and 27 parks are "cold spots" (low-value clusters -- disadvantaged/neglected areas). This
    gives direct guidance for urban green-space justice: focusing investment and renewal resources on the cold-spot
    (disadvantaged) clusters reduces green-space inequality. Gi* hot-spot analysis is the standard spatial method for
    finding "where park value/quality concentrates geographically" -- it maps green-space inequality.

====================================================================================

#114  DBSCAN (Map)
    file: 114_DBSCAN_park_kumeleri.xlsx
  >> SCENARIO (narration):
    We cluster park locations by density with DBSCAN, separating isolated parks (noise).
  >> VARIABLE SELECTION:
    - Coordinates: lat, lon  (Map tab)

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Number of clusters = 4   noise (isolated) = 8   n = 115
    Park locations clustered by density

>> COMMENTARY (narration):
    On the map we clustered parks' geographic locations by density with DBSCAN: four dense park groups were found, and
    8 parks were flagged as "isolated/noise" points belonging to no cluster. DBSCAN's strength is that it does not
    require the number of clusters in advance and can capture irregularly shaped clusters -- and automatically separates
    sparse/distant parks. The four clusters show parks concentrate in particular regions; the isolated parks may be the
    lone park in a green-space-deficit area (priority to protect/support). This spatial clustering is a practical tool
    for defining green-space service regions, finding access gaps and managing geographically similar parks together.
    With this we complete the full 114-analysis tour of the Landscape Architecture package.

====================================================================================

#115  Mixed-Design (Split-Plot) ANOVA
    file: 115_mixed_anova_thermal_comfort_index.xlsx
  >> SCENARIO (narration):
    We follow 40 sites measured at three season levels (spring, summer, autumn); each belongs to one of two design groups (naturalistic / hardscape). A mixed (split-plot) design tests the between-subjects
    main effect, the within-subjects main effect and their interaction on thermal comfort index.
  >> VARIABLE SELECTION:
    - Dependent variable: thermal_comfort_index
    - Subject ID: site_id
    - Between-subjects factor: design
    - Within-subjects factor: season

  ----- RESULT & COMMENTARY -----
>> RESULT (screen):
    Between (design): F(1,38) = 33.14  p < .001  np2 = 0.466
    Within (season):  F(2,76) = 19.42  p < .001  np2 = 0.338
    Interaction:         F(2,76) = 16.17  p < .001  np2 = 0.298
    Mauchly W = 0.915  p = 0.186   n_subjects = 40   n_obs = 120

>> COMMENTARY (narration):
    In a design x season mixed design we analyzed thermal comfort index for 40 sites (120 observations). The interaction is significant (F(2,76) = 16.17, p < .001, np2 = 0.298) *** -- the two groups' change across season differs in magnitude. The between-subjects main effect (naturalistic vs hardscape) is F = 33.14, p < .001; the within-subjects main effect (spring/summer/autumn) is F = 19.42, p < .001. Mauchly's test p = 0.186, so sphericity holds, so uncorrected within p is read directly. Read the interaction first: when it is significant the group effect must be interpreted separately at each season level. In landscape architecture, the mixed design is the standard analysis for comparing landscape designs for comfort across seasons.

====================================================================================
