Bayesian t-Test (BEST)

Bayesian t-Test (BEST)

⚡ Advanced · Bayesian Statistics · BEST

v1.0.2

Education Sciences context — the posterior-focused version of the classic t-test. With BF₁₀, it directly answers the question “how many times more evidence does the data provide in favor of H₁?” Three modes: one-sample/independent/paired. Rouder et al. (2009) JZS Cauchy prior.

BF₁₀ + Jeffreys interpretationJZS prior r=0.707Cohen’s d + %95 CrIPyMC GEREKMEZ

🎯 What is it for?

The Bayesian t-Test frees you from the classical p-value dilemma. BF₁₀ > 3: moderate-to-strong evidence for H₁; BF₁₀ < 1/3: moderate-to-strong evidence for H₀. In Education Sciences, it tests whether the Classical and Active Learning groups differ in terms of reading_comprehension_score.

📌 When is it used?

  • When comparing two groups (Classical vs. Active Learning) in Education Sciences and you want to test H₀ directly as well
  • Small sample (N<30) — the JZS prior stabilizes the estimates
  • In replication studies — seeking evidence for a “no effect” conclusion
  • Testing for a difference from a threshold (reference) value (one-sample mode)

⚙ Assumptions

  1. Continuous DV (reading_comprehension_score).
  2. Independence (independent mode): the observations of the two groups are independent.
  3. Normality is relaxed via the CLT; with small N, the Bayesian approach is more robust.
  4. A sensitivity analysis with JZS scale r ∈ {0.5, 0.707, 1.0} is recommended.

📊 How to Run in MerQur

1
Data tab → load the dataset.
2
Analysis → ⚡ Advanced → Bayesian t-Test (BEST).
3

Panel assignments (form fields in the program):

  • Columns: {'mode': 'independent', 'group_col': 'grup', 'val_col': 'birim_id'}
  • Parameters: {'prior_scale': 0.707}
4
JZS scale r: 0.707 (default). For sensitivity analysis you can also try 0.5 and 1.0.
5
▶ Run. BF₁₀ + Jeffreys interpretation scale chart + Cohen’s d.

📊 Sample Dataset — Education Sciences

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Education Sciences:

Egitim_Bilimleri/101_bayesian_t_test_new_old.xlsx

🎬 Scenario

Suppose a school district piloted a new instructional program and wants to
know whether it raises a student outcome compared with the old program. Each
of 80 learning units was assigned to either the old or the new program,
recorded in the group variable, and we measured the resulting outcome in
value. Instead of a classical t-test that only tells us whether to reject a
null hypothesis, we run a Bayesian t-test so we can quantify how much more
probable the difference is under the data and report a Bayes factor and a
credible interval for the effect. This is the right choice because we want a
continuous outcome compared across two independent groups while expressing
our evidence in probabilistic terms.

⚙️ Variable Selection

  • Grouping factor (two levels): group (old / new)
  • Dependent variable: value

Data Preview (First 5 Rows)

unit_id group value
1 old 60.02
2 old 54.28
3 old 67.56
4 old 35.73
5 old 42.32

n = 80 · Columns: unit_id, group, value

📈 MerQur Output

BAYESIAN T-TEST (BEST) RESULT
─────────────────────────────────────────────

BF10 = 3.36e+46 (decisive evidence) Cohen’s d = 3.94
New method vs old method — score

💬 Interpretation

We tested whether a new teaching method differs from the old with a Bayesian t-test. The Bayes Factor BF10 =
3.36e+46 — astronomically large, “decisive evidence”: the data support the hypothesis of a difference quadrillions
of times more than no difference. Cohen’s d = 3.94 makes the effect enormous. Unlike a classic p-value, the Bayes
Factor directly measures the strength of evidence for both H1 and H0 and distinguishes “absence of evidence” from
“evidence of absence”. The new method’s effect is not just statistical but practically overwhelming — the
Bayesian framework shows this convincingly.

⚠ Common Mistakes

  • In one-sample mode, leaving μ₀ at zero produces astronomical BF values for Likert/score comparisons; enter the correct reference point.
  • BF₁₀ > 3 and p < .05 do not always coincide — at small N, p may be significant while BF is weak.
  • Do not skip reporting prior sensitivity (r = 0.5/0.707/1.0).

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.