Bayesian t-Test (BEST)

Bayesian t-Test (BEST)

⚡ Advanced · Bayesian Statistics · BEST

v1.0.2

Architecture, Planning & Design context — the posterior-focused version of the classical t-test. With BF₁₀ it directly answers the question “how many times more evidence does the data provide in favor of H₁?” Three modes: one-sample/independent/paired. Rouder et al. (2009) JZS Cauchy prior.

BF₁₀ + Jeffreys interpretationJZS prior r=0.707Cohen’s d + %95 CrIPyMC GEREKMEZ

🎯 What is it for?

The Bayesian t-Test frees you from the classic p-value dilemma. BF₁₀ > 3: moderate-to-strong evidence for H₁; BF₁₀ < 1/3: moderate-to-strong evidence for H₀. In Architecture, Planning & Design it tests whether the Conventional and Green-Certified groups differ in enerji_tuketimi_kwh_m2.

📌 When is it used?

  • In Architecture, Planning & Design, when comparing two groups (Conventional vs Green-Certified) and you also want to test H₀ directly
  • Small samples (N<30) — the JZS prior stabilizes the estimates
  • In replication studies — seeking evidence for a “no effect” conclusion
  • Testing for a difference from a threshold (reference) value (one-sample mode)

⚙ Assumptions

  1. The DV is continuous (enerji_tuketimi_kwh_m2).
  2. Independence (independent mode): the observations of the two groups are independent.
  3. Normality is relaxed via the CLT; at small N the Bayesian approach is more robust.
  4. A sensitivity check with JZS scale r ∈ {0.5, 0.707, 1.0} is recommended.

📊 How Is It Run in MerQur?

1
Data tab → load the dataset.
2
Analysis → ⚡ Advanced → Bayesian t-Test (BEST).
3

Panel assignments (form fields in the program):

  • Columns: {'mode': 'independent', 'group_col': 'tasarim_grup', 'val_col': 'memnuniyet'}
  • Parameters: {'prior_scale': 0.707}
4
JZS scale r: 0.707 (default). For sensitivity analysis you may also try 0.5 and 1.0.
5
▶ Run. BF₁₀ + a Jeffreys interpretation-scale chart + Cohen’s d.

📊 Sample Dataset — Architecture, Planning & Design

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Architecture, Planning & Design:

Peyzaj_Mimarligi/101_bayesian_t_test_new_old_design.xlsx

🎬 Scenario

We evaluate the satisfaction difference of new vs old design with the Bayes Factor
(strength of evidence).

⚙️ Variable Selection

  • Dependent: satisfaction | Group: design_group

Data Preview (First 5 Rows)

park_id design_group satisfaction
1 old 3.53
2 old 3.91
3 old 3.31
4 old 3.45
5 old 3.24

n = 100 · Columns: park_id, design_group, satisfaction

📈 MerQur Output

BAYESIAN T-TEST (BEST) RESULT
─────────────────────────────────────────────

BF10 = 8.32e+67 (decisive evidence) Cohen’s d = 4.95
New design vs old design — satisfaction

💬 Interpretation

We tested whether a new design approach gives different satisfaction from the old with a Bayesian t-test. The Bayes
Factor BF10 = 8.32e+67 — astronomically large, “decisive evidence”: the data support the hypothesis of a
difference unimaginably more than no difference. Cohen’s d = 4.95 makes the effect extraordinarily large. The new
design raised satisfaction overwhelmingly. Unlike a classic p-value, the Bayes Factor directly measures the
strength of evidence for both H1 and H0 and distinguishes “absence of evidence” from “evidence of absence”. The new
design’s effect is not just statistical but practically overwhelming — the Bayesian framework shows this
convincingly.

⚠ Common Mistakes

  • In one-sample mode, leaving μ₀ at zero produces astronomical BF values for Likert/score comparisons; enter the correct reference point.
  • BF₁₀ > 3 and p < .05 do not always coincide — at small N, p may be significant while BF is weak.
  • Do not skip reporting prior sensitivity (r = 0.5/0.707/1.0).

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.