Hierarchical Bayesian Regression

Hierarchical Bayesian Regression

⚡ Advanced · Bayesian Statistics · Hierarchical

v1.0.2

Sport Sciences context — multilevel regression with a random intercept that varies across hakem. It estimates the puanlama_tutarlilik ~ deneyim_yili relationship separately for each referee and reports the posterior distribution + 95% CrI. PyMC NOT REQUIRED — statsmodels MixedLM + Empirical Bayes BLUP.

Multi-level BayesianRandom interceptBLUP posteriorBF₁₀ (LMM vs OLS)Caterpillar plot

🎯 What is it for?

If your data are hierarchical (student-school, patient-hospital, measurement-device, etc.), OLS regression ignores within-group correlation → type-I error increases. Hierarchical Bayesian learns each group’s own intercept and reports it together with the 95% CrI. With BF₁₀ (LMM vs OLS), whether group-level variance is needed is tested.

📌 When is it used?

  • Referee-level clustered data in Sport Sciences (≥10-20 individuals per hakem)
  • ICC > 0.05 — group-level variance should be taken into account
  • For the question “Which hakem shows high/low performance?”, shrinkage-corrected ranking

⚙ Assumptions

  1. Random effect ~ Normal(0, σ²hakem).
  2. Within-group residual ~ Normal(0, σ²).
  3. ≥10 observations per group are recommended; if fewer, an informative prior is critical.

📊 How to Run in MerQur

1
Load the data.
2
Analysis → ⚡ Advanced → Hierarchical Bayesian Regression.
3

Panel assignments (form fields in the program):

  • Columns: {'dv': 'Y_yanit', 'fixed': ['X_kovariat'], 'group': 'grup_id'}
  • Parameters: {}
4
REML estimator (default). EB BLUP automatic.
5
▶ Run. BF₁₀, ICC, caterpillar plot.

📊 Sample Dataset — Sport Sciences

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Sport Sciences:

Spor_Bilimleri/104_hierarchical_bayesian_LMM.xlsx

🎬 Scenario

We analyze group-nested data with a Bayesian hierarchical model. For stable
estimates in small groups, this is appropriate.

⚙️ Variable Selection

  • Dependent variable: Y_response
  • Cluster: group_id
  • Predictor(s): X_covariate

Data Preview (First 5 Rows)

group_id X_covariate Y_response
G1 48.3 50.08
G1 54.3 52.72
G1 56.66 63.58
G1 53.67 55.23
G1 54.54 55.56

n = 200 · Columns: group_id, X_covariate, Y_response

📈 MerQur Output

HIERARCHICAL BAYESIAN REGRESSION RESULT
─────────────────────────────────────────────

sigma^2_u (group) = 2.90 (8.6%) sigma^2_eps (residual) = 30.79 (91.4%) outcome: Y_response, group: group_id

💬 Interpretation

We analyzed group-nested data with a Bayesian hierarchical (multilevel) model: 8.6% of variability is
between-group, 91.4% within-group (ICC ~ 0.09). The Bayesian hierarchical model is the Bayesian version of LMM — it
estimates group effects with posterior distributions and balances small groups via “partial pooling”. In sports
science it is powerful for producing stable estimates even in small groups within multilevel (club/team/athlete)
data.

⚠ Common Mistakes

  • If there are <5 observations per group, the group-level estimate is unreliable — a more informative prior is required.
  • A random slope (the slope of random deneyim_yili) is not the default; if there is a cross-level interaction, it must be added explicitly.
  • If ICC < 0.05, OLS may be sufficient — check BF₁₀.

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.