K-Means Clustering
K-Means Clustering is one of the statistical analyses applied automatically in MerQur. This page illustrates — on a Education Sciences sample dataset — how the analysis is run, what the MerQur output looks like, and how the result is reported in APA 7 format.
🎯 What is it for?
K-Means Clustering automatically applies in the background all the assumption checks required for the relevant data type (normality, homogeneity of variance, etc.) and presents the results with a clear table + chart. Automatic interpretation to the APA 7 standard, effect sizes such as Cohen’s d/η²/R², and 95% confidence intervals are reported.
📌 When is it used?
- Statistical analysis of measurements in the Education Sciences domain
- To produce APA 7-compatible result tables for academic publications
- Hypothesis testing and decision-making processes
- Undergraduate / master’s / PhD theses after the appropriate method has been selected
📐 Assumptions
- Appropriate scale — Variables must be at the measurement level required by the analysis (nominal/ordinal/interval/ratio)
- Independent observations — Observations must come from individuals independent of one another
- Sufficient sample size — The minimum n requirement for the analysis must be met
- Outlier check — Outliers must be detected and evaluated
If assumptions are violated, MerQur automatically suggests a non-parametric or robust alternative.
🛠 How to do it in MerQur
Load the data. Select the sample file from File → Open. MerQur auto-detects column types.
Select the analysis. From the left side select K-Means Clustering.
Panel assignments (form fields in the program):
- Columns:
['ogrenci_id', 'GPA', 'calisma_saat', 'matematik', 'deneme_puan'] - Clusters:
3 - init:
k-means++ - Missing strategy:
drop
Optional settings. Effect size ✓ · 95% confidence interval ✓ · Assumption checks (automatic).
▶ Run — click the button. Results are produced automatically as a table + chart.
📄 Export to Word. APA 7-formatted report with italic statistical symbols.
📊 Sample Dataset — Education Sciences
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Education Sciences:
Egitim_Bilimleri/66_kmeans_student_4kume.xlsx
🎬 Scenario
Let’s say we want to discover natural groupings of students based on their academic behavior, without any predefined labels. For 240 students we have GPA, weekly study_hour, math score, and a trial_score. K-Means Clustering is the right tool because it partitions students into a chosen number of homogeneous clusters by minimizing within-group distance, helping us identify profiles such as high-achieving hard workers or low-engagement learners purely from the data.
⚙️ Variable Selection
- Clustering variables: GPA
- Clustering variables: study_hour
- Clustering variables: math
- Clustering variables: trial_score
Data Preview (First 5 Rows)
| student_id | GPA | study_hour | math | trial_score |
|---|---|---|---|---|
| 1.0 | 2.14 | 3.8 | 49.6 | 53.5 |
| 2.0 | 2.84 | 14.0 | 83.7 | 74.0 |
| 3.0 | 3.73 | 22.1 | 81.3 | 100.0 |
| 4.0 | 3.45 | 36.5 | 90.5 | 95.4 |
| 5.0 | 2.21 | 2.0 | 33.4 | 45.7 |
n = 240 · Columns: student_id, GPA, study_hour, math, trial_score
📈 MerQur Output
💬 Interpretation
We split students into four natural groups (a typology) by their academic features with K-Means. Silhouette = 0.46 is moderate — the clusters separate but not sharply, which is expected in educational data (achievement is a continuous gradient, sharp breaks are rare). The four groups likely represent distinct student profiles: high-achieving hard-workers, low-achieving low-effort, and intermediate types. K-Means finds natural groups in unlabeled data; it is practical for student profiling, designing targeted interventions and identifying risk groups.
⚠ Common Mistakes
- Misidentifying the data type (e.g., loading a categorical variable as numeric)
- Skipping assumption checks and going straight to the p-value
- Failing to report effect size — APA 7 requires both p and effect size
- Failing to apply a Type I error correction (Bonferroni/Tukey) in multiple comparisons
- Not switching to a non-parametric alternative when n is insufficient
📹 Video Walkthrough
Watch the video below for an end-to-end walkthrough of this analysis on a Education Sciences file.
▶ K-Means Clustering — video walkthrough
This analysis is demonstrated on a sample dataset from Anaesthesiology (the steps are identical across disciplines). The link jumps straight to 0:00. Narration is in Turkish.
🧪 Natural Sciences & Mathematics · 🏛 Architecture, Planning & Design · ⚙ Engineering · 🏥 Health Sciences · 📊 Social, Humanities & Admin Sciences · 🏃 Sport Sciences · 🌾 Agriculture, Forestry & Aquatic
📚 If You Used This Analysis, Cite MerQur
If you performed this analysis using MerQur in a scientific study, please use the citation below as part of your academic citation obligations (APA 7):
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
For BibTeX, RIS, EndNote and the English citation form: all citation formats →
- American Psychological Association. (2020). Publication manual of the American Psychological Association (7th ed.).
- Field, A. (2018). Discovering statistics using IBM SPSS Statistics (5th ed.). Sage.
- Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum.