Bayesian t-Test (BEST)

Bayesian t-Test (BEST)

⚡ Advanced · Bayesian Statistics · BEST

v1.0.2

Health Sciences context — a posterior-focused version of the classical t-test. With BF₁₀, it gives a direct answer to the question “how many times more evidence does the data provide in favor of H₁?” Three modes: one-sample/independent/paired. Rouder et al. (2009) JZS Cauchy prior.

BF₁₀ + Jeffreys interpretationJZS prior r=0.707Cohen’s d + %95 CrIPyMC GEREKMEZ

🎯 What is it for?

The Bayesian t-Test frees you from the classical p-value dilemma. BF₁₀ > 3: moderate-to-strong evidence for H₁; BF₁₀ < 1/3: moderate-to-strong evidence for H₀. In the Health Sciences it tests whether the Placebo and Drug-X groups differ in terms of sistolik_kan_basinci.

📌 When is it used?

  • In the Health Sciences, when comparing two groups (Placebo vs Drug-X) and you also want to test H₀ directly
  • Small samples (N<30) — the JZS prior stabilizes the estimates
  • In replication studies — seeking evidence for a “no effect” finding
  • Testing a difference from a threshold (reference) value (one-sample mode)

⚙ Assumptions

  1. The DV is continuous (sistolik_kan_basinci).
  2. Independence (independent mode): the observations of the two groups are independent.
  3. Normality is relaxed via the CLT; with small N, the Bayesian approach is more robust.
  4. A sensitivity test with JZS scale r ∈ {0.5, 0.707, 1.0} is recommended.

📊 How to Run It in MerQur

1
Data tab → load the data set.
2
Analysis → ⚡ Advanced → Bayesian t-Test (BEST).
3

Panel assignments (form fields in the program):

  • Columns: {'mode': 'independent', 'group_col': 'tedavi', 'val_col': 'yanit_skoru'}
  • Parameters: {'prior_scale': 0.707}
4
JZS scale r: 0.707 (default). For sensitivity you may also try 0.5 and 1.0.
5
▶ Run. BF₁₀ + Jeffreys interpretation scale chart + Cohen’s d.

📊 Sample Dataset — Health Sciences

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Health Sciences:

Tip/101_bayesian_t_test_new_old_treatment.xlsx

🎬 Scenario

We examine two treatment groups’ response-score difference with a Bayesian
t-test, expressing evidence as a Bayes factor. For an intuitive evidence ratio,
this is appropriate.

⚙️ Variable Selection

  • Dependent variable: response_score
  • Grouping (categorical): treatment

Data Preview (First 5 Rows)

patient_id treatment response_score
1 old 64.6
2 old 72.18
3 old 60.18
4 old 62.92
5 old 58.74

n = 70 · Columns: patient_id, treatment, response_score

📈 MerQur Output

BAYESIAN T-TEST (BEST) RESULT
─────────────────────────────────────────────

BF10 = 2.81e+40 Cohen’s d = 3.937 (overwhelming evidence for H1) outcome: response_score, group: treatment

💬 Interpretation

We examined the response-score difference between two treatment groups with a Bayesian t-test: the Bayes factor is
overwhelming (BF10 ~ 2.8e40), the effect very large (d = 3.94) — the data support the “difference” hypothesis over
“no difference” by astronomical odds. Unlike a p-value, the Bayes factor gives the RELATIVE evidence strength of two
hypotheses and can distinguish “no evidence” from “no difference”. In medicine it is preferred when one wants to
express the evidential strength of a treatment decision as an intuitive ratio.

⚠ Common Mistakes

  • In one-sample mode, leaving μ₀ at zero produces astronomical BF values for Likert/score comparisons; enter the correct reference point.
  • BF₁₀ > 3 and p < .05 do not always coincide — at small N, p may be significant while BF is weak.
  • Do not skip reporting prior sensitivity (r = 0.5/0.707/1.0).

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.