Bayesian t-Test (BEST)

Bayesian t-Test (BEST)

⚡ Advanced · Bayesian Statistics · BEST

v1.0.2

Agriculture, Forestry & Aquatic context — a posterior-focused version of the classic t-test. With BF₁₀, it directly answers the question “how many times more evidence does the data provide in favor of H₁?” Three modes: one-sample/independent/paired. Rouder et al. (2009) JZS Cauchy prior.

BF₁₀ + Jeffreys interpretationJZS prior r=0.707Cohen’s d + %95 CrIPyMC GEREKMEZ

🎯 What is it for?

The Bayesian t-Test frees you from the classic p-value dilemma. BF₁₀ > 3: moderate-to-strong evidence for H₁; BF₁₀ < 1/3: moderate-to-strong evidence for H₀. In Agriculture, Forestry & Aquatic it tests whether the Conventional and Organic groups differ in terms of urun_verimi_kg_da.

📌 When is it used?

  • When comparing two groups (Conventional vs Organic) in Agriculture, Forestry & Aquatic and you also want to test H₀ directly
  • Small samples (N<30) — the JZS prior stabilizes the estimates
  • In replication studies — searching for evidence of a “no effect” result
  • Testing a difference from a threshold (reference) value (one-sample mode)

⚙ Assumptions

  1. The DV is continuous (urun_verimi_kg_da).
  2. Independence (independent mode): the observations of the two groups are independent.
  3. Normality is relaxed via the CLT; for small N the Bayesian approach is more robust.
  4. A sensitivity check with JZS scale r ∈ {0.5, 0.707, 1.0} is recommended.

📊 How to Run in MerQur

1
Data tab → load the dataset.
2
Analysis → ⚡ Advanced → Bayesian t-Test (BEST).
3

Panel assignments (form fields in the program):

  • Columns: {'mode': 'independent', 'group_col': 'grup', 'val_col': 'birim_id'}
  • Parameters: {'prior_scale': 0.707}
4
JZS scale r: 0.707 (default). For sensitivity analysis you can also try 0.5 and 1.0.
5
▶ Run. BF₁₀ + Jeffreys interpretation scale chart + Cohen’s d.

📊 Sample Dataset — Agriculture, Forestry & Aquatic

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Agriculture, Forestry & Aquatic:

Ziraat_Orman_Su/101_bayesian_t_test_new_old.xlsx

🎬 Scenario

We test the difference between two applications, new and old (group), on a measurement (value) from a Bayesian framework. The
classical t-test gives only a p-value; the Bayesian t-test, with the Bayes factor (BF10), answers “by how many times do the
data support one hypothesis over the other” and shows the posterior distribution of the effect size: how strong is the evidence
for a difference?

⚙️ Variable Selection

  • Dependent (continuous): value
  • Grouping (2 categories): group

Data Preview (First 5 Rows)

unit_id group value
1 old 60.02
2 old 54.28
3 old 67.56
4 old 35.73
5 old 42.32

n = 80 · Columns: unit_id, group, value

📈 MerQur Output

BAYESIAN T-TEST (BEST) RESULT
─────────────────────────────────────────────

BF10 = 3.4 x 10^46 (decisive evidence, H1) Cohen’s d = 3.94 (value ~ group: new/old)

💬 Interpretation

We tested the difference between a new and an old application on a measurement from a Bayesian framework. The
classical t-test gives only a p-value; the Bayesian t-test, with the Bayes factor, answers “by how many times do
the data support one hypothesis over the other”. The result is staggering: BF10 = 3.4 x 10 to the 46. This means
the data support “a difference exists” over “no difference” by an astronomical ratio — far beyond “decisive
evidence”. Cohen’s d = 3.94 means the effect is enormous too. The beauty of the Bayes factor is doing what the
p-value cannot: it can also measure evidence FOR H0 and gives the STRENGTH of evidence on a continuous scale. Here
the difference is so clear that we can say the new application is decisively superior to the old.

⚠ Common Mistakes

  • In one-sample mode, leaving μ₀ at zero produces astronomical BF values for Likert/score comparisons; enter the correct reference point.
  • BF₁₀ > 3 and p < .05 do not always coincide — at small N, p may be significant while BF is weak.
  • Do not skip reporting prior sensitivity (r = 0.5/0.707/1.0).

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.