Spatial Error Model

Spatial Error Model

⚡ Advanced · Spatial Regression

v1.0.2

Health Sciences context — Spatial Error Model. A regression that accounts for the spatial neighborhood structure of diyabet_prevalans_pct. Spatial weight matrix: K-nearest neighbor (k=5), row-standardized.

Y = Xβ + λWu + εlibpysal KNN Wspreg MLλ ≈ 0.62

🎯 What is it for?

Classic OLS assumes the observations are independent — but in lat/lon data, nearby points take on similar values (spatial autocorrelation). The Spatial Error Model incorporates this structure into the model; otherwise the standard errors of the OLS β estimates come out too small (type-I error).

📌 When is it used?

  • Spatial data (with lat/lon) + a continuous DV in the Health Sciences
  • Moran’s I p < .05 — evidence of spatial autocorrelation
  • Spatial clustering in OLS residuals (hotspot over DBSCAN)

⚙ Assumptions

  1. Lat/lon coordinates ({lat, lon}).
  2. Continuous DV (diyabet_prevalans_pct).
  3. Spatial weight matrix design (KNN k=5 — default, with dist.band as an alternative).
  4. The λ parameter is stable within 0-1; there is a risk of fragility at the boundary.

📊 How to Run It in MerQur

1
Load the data (lat, lon, DV, X1, X2 columns).
2
Analysis → ⚡ Advanced → Spatial Error Model.
3

Panel assignments (form fields in the program):

  • Columns: {'y': 'vaka_orani', 'x': ['sosyoekonomik', 'yas_ortalama'], 'lat': 'lat', 'lon': 'lon'}
  • Parameters: {'weights': 'knn', 'k': 5}
4
Estimator: ML (default).
5
▶ Run. Coefficient forest plot + λ + map.

📊 Sample Dataset — Health Sciences

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Health Sciences:

Tip/106_spatial_error_disease_residual_autocorr.xlsx

🎬 Scenario

We model spatial dependency in the error term. For unmeasured geographic
factors, SEM is appropriate.

⚙️ Variable Selection

  • Dependent variable: case_ratio
  • Predictor(s): socioeconomic
  • Predictor(s): age_mean
  • Latitude: lat
  • Longitude: lon

Data Preview (First 5 Rows)

province_id lat lon socioeconomic age_mean case_ratio
1.0 41.378 40.0929 33.5 34.0 501.3
2.0 37.9454 38.0209 59.6 42.9 632.8
3.0 40.0527 27.3987 23.5 44.2 491.1
4.0 37.8009 33.2798 73.1 50.3 547.6
5.0 38.4268 31.0605 56.5 36.5 476.4

n = 100 · Columns: province_id, lat, lon, socioeconomic, age_mean, case_ratio

📈 MerQur Output

SPATIAL ERROR MODEL RESULT
─────────────────────────────────────────────

lambda = 0.630 z = 7.08 p < .001 *** Pseudo R^2 = 0.72 (N = 100, k-NN W, k = 5)

💬 Interpretation

This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
0.63, p < .001) — the effect of geographic variables omitted from the model makes neighboring errors correlated.
Unlike SAR, SEM attributes the spread to the error rather than the outcome. In medicine/epidemiology, when the source
of spatial autocorrelation is unmeasured geographic factors (health access, environment), the correct specification
is SEM; it is chosen by comparison with SAR.

⚠ Common Mistakes

  • λ ≈ 1 → matrix singularity; try a looser W (k=10).
  • The choice of W matrix affects the results — compare KNN and DistanceBand.
  • Projected coordinates (UTM) instead of Lat/Lon may be preferable for accurate distance.

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.