Spatial Error Model
v1.0.2
Health Sciences context — Spatial Error Model. A regression that accounts for the spatial neighborhood structure of diyabet_prevalans_pct. Spatial weight matrix: K-nearest neighbor (k=5), row-standardized.
🎯 What is it for?
Classic OLS assumes the observations are independent — but in lat/lon data, nearby points take on similar values (spatial autocorrelation). The Spatial Error Model incorporates this structure into the model; otherwise the standard errors of the OLS β estimates come out too small (type-I error).
📌 When is it used?
- Spatial data (with lat/lon) + a continuous DV in the Health Sciences
- Moran’s I p < .05 — evidence of spatial autocorrelation
- Spatial clustering in OLS residuals (hotspot over DBSCAN)
⚙ Assumptions
- Lat/lon coordinates ({lat, lon}).
- Continuous DV (diyabet_prevalans_pct).
- Spatial weight matrix design (KNN k=5 — default, with dist.band as an alternative).
- The λ parameter is stable within 0-1; there is a risk of fragility at the boundary.
📊 How to Run It in MerQur
Panel assignments (form fields in the program):
- Columns:
{'y': 'vaka_orani', 'x': ['sosyoekonomik', 'yas_ortalama'], 'lat': 'lat', 'lon': 'lon'} - Parameters:
{'weights': 'knn', 'k': 5}
📊 Sample Dataset — Health Sciences
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Health Sciences:
Tip/106_spatial_error_disease_residual_autocorr.xlsx
🎬 Scenario
We model spatial dependency in the error term. For unmeasured geographic
factors, SEM is appropriate.
⚙️ Variable Selection
- Dependent variable: case_ratio
- Predictor(s): socioeconomic
- Predictor(s): age_mean
- Latitude: lat
- Longitude: lon
Data Preview (First 5 Rows)
| province_id | lat | lon | socioeconomic | age_mean | case_ratio |
|---|---|---|---|---|---|
| 1.0 | 41.378 | 40.0929 | 33.5 | 34.0 | 501.3 |
| 2.0 | 37.9454 | 38.0209 | 59.6 | 42.9 | 632.8 |
| 3.0 | 40.0527 | 27.3987 | 23.5 | 44.2 | 491.1 |
| 4.0 | 37.8009 | 33.2798 | 73.1 | 50.3 | 547.6 |
| 5.0 | 38.4268 | 31.0605 | 56.5 | 36.5 | 476.4 |
n = 100 · Columns: province_id, lat, lon, socioeconomic, age_mean, case_ratio
📈 MerQur Output
─────────────────────────────────────────────
lambda = 0.630 z = 7.08 p < .001 *** Pseudo R^2 = 0.72 (N = 100, k-NN W, k = 5)
💬 Interpretation
This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
0.63, p < .001) — the effect of geographic variables omitted from the model makes neighboring errors correlated.
Unlike SAR, SEM attributes the spread to the error rather than the outcome. In medicine/epidemiology, when the source
of spatial autocorrelation is unmeasured geographic factors (health access, environment), the correct specification
is SEM; it is chosen by comparison with SAR.
⚠ Common Mistakes
- λ ≈ 1 → matrix singularity; try a looser W (k=10).
- The choice of W matrix affects the results — compare KNN and DistanceBand.
- Projected coordinates (UTM) instead of Lat/Lon may be preferable for accurate distance.
📚 MerQur’a Atıf
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.