VARCOMP (Variance-Components Analysis)
v1.0.1
Education Sciences context — in a hierarchical design, it decomposes how much of the total variance originates from each level (site, group, sub-group, error). Equivalent to SAS PROC VARCOMP, implemented in MerQur with a REML/ML cascade.
REML + ML fallbackICC + heritabilityσ²okul ≈ 8.0
🎯 What is it for?
VARCOMP answers the question of which level contributes how much variance. Is the variability in the distribution of a single akademik_basari_puan measurement entirely random, or does it stem from systematic differences between okul and sinif? With VARCOMP, σ²okul, σ²sinif, σ²ogrenci and σ²residual are estimated separately; the percentage share of each is reported.
📌 When is it used?
- When there is a multilevel sampling design (nested) in the Education Sciences field
- For sample-size calculation when designing a data collection plan in which you will take the same measurement on different sinif’s under different okul’s
- To compute heritability / generalizability (σ² ratios)
- Before even building the fixed-effect model — where does the variance primarily lie?
⚙ Assumptions
- Continuous DV (akademik_basari_puan).
- Nested structure: sinif within okul, ogrenci within sinif (no cross-level confounding).
- Random effects ~ Normal (homoscedastic, mean 0).
- Sufficient units at each level: ≥4-6 okul, with ≥2 sinif in each recommended.
📊 How to Run It in MerQur?
Panel assignments (form fields in the program):
- Columns:
{'dv': 'deger', 'factors': [('ust_birim', False), ('alt_birim', True)]} - Parameters:
{'estimator': 'REML'}
📊 Sample Dataset — Education Sciences
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Education Sciences:
Egitim_Bilimleri/100_varcomp_3seviye_h2.xlsx
🎬 Scenario
Imagine we want to understand where the variability in student achievement
scores actually comes from in a three-level school system. Our 100
measurements are nested: each value belongs to a lower_unit (such as a
class), which is nested within an upper_unit (one of four districts, L_A to
L_D), with a measurement_id identifying the individual record. Rather than
testing a single mean difference, we want to know what proportion of the
total variance lies between districts, between classes, and within classes.
We use Variance Components Analysis to partition the variance of value
across these nested levels and estimate, for example, an intraclass measure
of how much districts and classes matter.
⚙️ Variable Selection
- Dependent variable (value): value
- Random level 1 (upper): upper_unit (L_A/L_B/L_C/L_D)
- Random level 2 (lower, nested): lower_unit
- Measurement / observation id: measurement_id
Data Preview (First 5 Rows)
| upper_unit | lower_unit | measurement_id | value |
|---|---|---|---|
| L_A | L_A_S01 | L_A_S01_M01 | 51.41 |
| L_A | L_A_S01 | L_A_S01_M02 | 52.17 |
| L_A | L_A_S01 | L_A_S01_M03 | 40.6 |
| L_A | L_A_S01 | L_A_S01_M04 | 43.19 |
| L_A | L_A_S01 | L_A_S01_M05 | 48.91 |
n = 100 · Columns: upper_unit, lower_unit, measurement_id, value
📈 MerQur Output
─────────────────────────────────────────────
upper unit = 12.1% lower unit (nested in upper) = 40.9% residual = 47.0%
Nested variance components (REML)
💬 Interpretation
We partitioned the variability of a measurement by which level — upper unit (e.g. school), lower unit (class) or
individual — it comes from with variance-components analysis. The result: 40.9% of the variance is at the lower
unit (class) level, 12.1% at the upper unit (school) level, and 47% individual/residual variation. So the largest
structural source is the class level; between-school differences are relatively small. This is a critical
inference in heritability (h2) and multi-level design studies: it shows where the variability concentrates and at
which level (class, school or individual) an intervention should be targeted.
⚠ Common Mistakes
- Forgetting the nested checkbox — the sinif column is clustered under okul (e.g. SINIF_05_03 exists only in OKUL_05). If the checkbox is not checked, a crossed structure is assumed → variance is misallocated.
- A single level at the second tier — if there is only one sinif under an okul, VARCOMP cannot estimate that level (σ²=sinif≈0 ‘singular fit’ warning).
- Continuous DV assumption — for Bernoulli/count data, a GLMM or a categorical-VARCOMP alternative is required instead of VARCOMP (outside MerQur’s scope).
📚 MerQur’a Atıf
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.