UMAP
UMAP is one of the statistical analyses applied automatically in MerQur. This page illustrates — on a Education Sciences sample dataset — how the analysis is run, what the MerQur output looks like, and how the result is reported in APA 7 format.
🎯 What is it for?
UMAP automatically applies in the background all the assumption checks required for the relevant data type (normality, homogeneity of variance, etc.) and presents the results with a clear table + chart. Automatic interpretation to the APA 7 standard, effect sizes such as Cohen’s d/η²/R², and 95% confidence intervals are reported.
📌 When is it used?
- Statistical analysis of measurements in the Education Sciences domain
- To produce APA 7-compatible result tables for academic publications
- Hypothesis testing and decision-making processes
- Undergraduate / master’s / PhD theses after the appropriate method has been selected
📐 Assumptions
- Appropriate scale — Variables must be at the measurement level required by the analysis (nominal/ordinal/interval/ratio)
- Independent observations — Observations must come from individuals independent of one another
- Sufficient sample size — The minimum n requirement for the analysis must be met
- Outlier check — Outliers must be detected and evaluated
If assumptions are violated, MerQur automatically suggests a non-parametric or robust alternative.
🛠 How to do it in MerQur
Load the data. Select the sample file from File → Open. MerQur auto-detects column types.
Select the analysis. From the left side select UMAP.
Panel assignments (form fields in the program):
- feature_cols:
[ogrenci_id, davranis_01, davranis_02, +23 daha] - n_components:
2 - n_neighbors:
15 - min_dist:
0.1
Optional settings. Effect size ✓ · 95% confidence interval ✓ · Assumption checks (automatic).
▶ Run — click the button. Results are produced automatically as a table + chart.
📄 Export to Word. APA 7-formatted report with italic statistical symbols.
📊 Sample Dataset — Education Sciences
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Education Sciences:
Egitim_Bilimleri/72_umap_behavior_5tip.xlsx
🎬 Scenario
Here we have 200 students rated on 25 behavioral indicators, and each student also belongs to one of five latent behavior profiles, labeled S1 through S5 in lower_type. We want to see whether these 25 behaviors naturally separate the students into the five known profile groups when projected onto a 2D plane. UMAP is ideal for this because it is a nonlinear dimensionality reduction technique that preserves local neighborhood structure, so behaviorally similar students stay together, and we can color the embedding by lower_type to check whether the profiles form distinct clusters.
⚙️ Variable Selection
- Features (high-dim input): behavior_01 … behavior_25 (all 25 behavior items)
- Color / Label (grouping): lower_type (S1, S2, S3, S4, S5)
Data Preview (First 5 Rows)
| student_id | lower_type | behavior_01 | behavior_02 | behavior_03 | behavior_04 | behavior_05 | behavior_06 | behavior_07 | behavior_08 | behavior_09 | behavior_10 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | S1 | 0.919 | -0.374 | -0.135 | 0.982 | 1.353 | -1.138 | 0.836 | -2.118 | 1.164 | -0.049 |
| 2 | S1 | -0.683 | 0.188 | -0.839 | 0.111 | -0.686 | 0.293 | 0.67 | -0.268 | -0.837 | 1.092 |
| 3 | S1 | -0.968 | -0.615 | -1.128 | 0.985 | 1.474 | -0.519 | -0.212 | 1.055 | 0.111 | -0.698 |
| 4 | S1 | -0.49 | -0.734 | -1.322 | -0.678 | -0.051 | 0.377 | 0.077 | -1.635 | 0.433 | -2.254 |
| 5 | S1 | -1.885 | 0.715 | 2.111 | 0.388 | 0.447 | 0.592 | 1.071 | 2.284 | -1.765 | -1.25 |
n = 200 · Columns (first 12 columns): student_id, lower_type, behavior_01, behavior_02, behavior_03, behavior_04, behavior_05, behavior_06, behavior_07, behavior_08, behavior_09, behavior_10
📈 MerQur Output
💬 Interpretation
UMAP is a modern dimension-reduction method similar to t-SNE but it preserves both local and global structure better and is faster. We projected student profiles of 25 behavior items into a two-dimensional map; similar behavior profiles cluster. The n_neighbors parameter tunes the local-global balance, and min_dist the tightness of clusters. UMAP is increasingly preferred for discovering hidden group structure in high-dimensional behavior/survey data. It is used for visualisation; the clustering pattern is interpreted, not the axis values.
⚠ Common Mistakes
- Misidentifying the data type (e.g., loading a categorical variable as numeric)
- Skipping assumption checks and going straight to the p-value
- Failing to report effect size — APA 7 requires both p and effect size
- Failing to apply a Type I error correction (Bonferroni/Tukey) in multiple comparisons
- Not switching to a non-parametric alternative when n is insufficient
📹 Video Walkthrough
Watch the video below for an end-to-end walkthrough of this analysis on a Education Sciences file.
▶ UMAP — video walkthrough
This section is part of the Kümeleme ve Boyut İndirgeme video (4 analyses in one video). The link below jumps straight to 5:22, where this analysis begins. Narration is in Turkish.
🧪 Natural Sciences & Mathematics · 🏛 Architecture, Planning & Design · ⚙ Engineering · 🏥 Health Sciences · 📊 Social, Humanities & Admin Sciences · 🏃 Sport Sciences · 🌾 Agriculture, Forestry & Aquatic
📚 If You Used This Analysis, Cite MerQur
If you performed this analysis using MerQur in a scientific study, please use the citation below as part of your academic citation obligations (APA 7):
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
For BibTeX, RIS, EndNote and the English citation form: all citation formats →
- American Psychological Association. (2020). Publication manual of the American Psychological Association (7th ed.).
- Field, A. (2018). Discovering statistics using IBM SPSS Statistics (5th ed.). Sage.
- Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum.