VARCOMP (Variance-Components Analysis)
v1.0.1
Agriculture, Forestry & Aquatic context — in a hierarchical design, decomposes how much of the total variance originates from each level (site, group, subgroup, error). Equivalent to SAS PROC VARCOMP, implemented in MerQur with a REML/ML cascade.
REML + ML fallbackICC + heritabilityσ²site ≈ 0.5
🎯 What is it for?
VARCOMP answers the question of how much variance contribution each level makes. Is the variability in the distribution of a single boy_artisi_m measurement entirely random, or does it stem from systematic differences between site and population? With VARCOMP, σ²site, σ²population, σ²family and σ²residual are estimated separately; the percentage share of each is reported.
📌 When is it used?
- When there is a multilevel (nested) sampling design in the Agriculture, Forestry & Aquatic domain
- For sample-size calculation when designing a data-collection plan in which you will take the same measurement on different populations under different sites
- To compute heritability / generalizability (σ² ratios)
- Before building a fixed-effects model at all — where is the variance primarily?
⚙ Assumptions
- Continuous DV (boy_artisi_m).
- Nested structure: population is within site, family is within population (no cross-level confounding).
- Random effects ~ Normal (homoskedastic, mean 0).
- Sufficient units at each level: ≥4-6 sites, each with ≥2 populations recommended.
📊 How to Run It in MerQur?
Panel assignments (form fields in the program):
- Columns:
{'dv': 'deger', 'factors': [('ust_birim', False), ('alt_birim', True)]} - Parameters:
{'estimator': 'REML'}
📊 Sample Dataset — Agriculture, Forestry & Aquatic
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Agriculture, Forestry & Aquatic:
Ziraat_Orman_Su/100_varcomp_3seviye_h2.xlsx
🎬 Scenario
In a three-level nested trial (upper_unit / lower_unit / measurement) we want to decompose from which level the total variance
of a trait (value) originates. Variance-components analysis partitions the variance into levels; in forestry/breeding this is
the basis for heritability (h2) estimation: how much of the variability in the trait is genetic/upper-group, and how much is
environmental/measurement?
⚙️ Variable Selection
- Dependent (continuous): value
- Upper level: upper_unit
- Lower level (nested): lower_unit
Data Preview (First 5 Rows)
| upper_unit | lower_unit | measurement_id | value |
|---|---|---|---|
| L_A | L_A_S01 | L_A_S01_M01 | 51.41 |
| L_A | L_A_S01 | L_A_S01_M02 | 52.17 |
| L_A | L_A_S01 | L_A_S01_M03 | 40.6 |
| L_A | L_A_S01 | L_A_S01_M04 | 43.19 |
| L_A | L_A_S01 | L_A_S01_M05 | 48.91 |
n = 100 · Columns: upper_unit, lower_unit, measurement_id, value
📈 MerQur Output
─────────────────────────────────────────────
Variance components (REML): upper_unit 12.1% | lower_unit (nested) 40.9% | residual 47.0%
dv = value
💬 Interpretation
In a nested trial we decomposed from which level the total variance of a trait originates. Variance-components
analysis, unlike LMM, uses no fixed effects — all factors are random, the goal being to partition variance into
levels. The result is instructive: 12 percent of the variability comes from the upper level (e.g. population), 41
percent from the lower level (e.g. family, nested within population), and 47 percent from the residual
(environment/measurement). In forestry and breeding this is the basis of heritability (h2) estimation: how much of
the variability in a trait is genetic/structural and how much environmental? The lower level (family) contributing
more than the upper (population) directly shows which level to focus on in selection/breeding work.
⚠ Common Mistakes
- Forgetting the nested checkbox — the population column is clustered under site (e.g. POPULASYON_05_03 exists only in SAHA_05). If the checkbox is not checked, a crossed structure is assumed → variance is partitioned incorrectly.
- A single level at the second tier — if there is only one population under a site, VARCOMP cannot estimate that level (σ²=population≈0 ‘singular fit’ warning).
- Continuous DV assumption — for Bernoulli/count data, a GLMM or a categorical-VARCOMP alternative is needed instead of VARCOMP (outside MerQur’s scope).
📚 MerQur’a Atıf
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.