Spatial Error Model

Spatial Error Model

⚡ Advanced · Spatial Regression

v1.0.2

Sport Sciences context — Spatial Error Model. A regression that accounts for the spatial neighborhood structure of il_spor_katilim_orani. Spatial weight matrix: K-nearest neighbor (k=5), row standardized.

Y = Xβ + λWu + εlibpysal KNN Wspreg MLλ ≈ 0.48

🎯 What is it for?

Classic OLS assumes the observations are independent — but in lat/lon data, nearby points take on similar values (spatial autocorrelation). The Spatial Error Model incorporates this structure into the model; otherwise the standard errors of the OLS β estimates come out too small (type-I error).

📌 When is it used?

  • In Sport Sciences, spatial data (with lat/lon) + continuous DV
  • Moran’s I p < .05 — evidence of spatial autocorrelation
  • Spatial clustering in OLS residuals (hotspot over DBSCAN)

⚙ Assumptions

  1. Lat/lon coordinates ({lat, lon}).
  2. Continuous DV (il_spor_katilim_orani).
  3. Spatial weight matrix design (KNN k=5 — default, dist.band alternative available).
  4. The λ parameter is stable between 0 and 1; risk of fragility at the boundary.

📊 How to Run in MerQur

1
Load the data (lat, lon, DV, X1, X2 columns).
2
Analysis → ⚡ Advanced → Spatial Error Model.
3

Panel assignments (form fields in the program):

  • Columns: {'y': 'birim_id', 'x': ['X1', 'X2'], 'lat': 'lat', 'lon': 'lon'}
  • Parameters: {'weights': 'knn', 'k': 5}
4
Estimator: ML (default).
5
▶ Run. Coefficient forest plot + λ + map.

📊 Sample Dataset — Sport Sciences

ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.

🎬 Example File

This analysis is demonstrated on the following example dataset for Sport Sciences:

Spor_Bilimleri/106_spatial_error_residual.xlsx

🎬 Scenario

We model spatial dependency in the error term. For unmeasured geographic
factors, SEM is appropriate.

⚙️ Variable Selection

  • Dependent variable: Y_value
  • Predictor(s): X1
  • Predictor(s): X2
  • Latitude: lat
  • Longitude: lon

Data Preview (First 5 Rows)

unit_id lat lon X1 X2 Y_value
1.0 40.729 36.2513 58.0 18.3 369.6
2.0 37.3016 39.779 60.9 46.8 460.7
3.0 38.3563 33.0728 73.2 48.6 495.5
4.0 41.2077 28.8098 82.7 15.0 199.5
5.0 40.7324 29.6924 48.9 32.5 410.3

n = 100 · Columns: unit_id, lat, lon, X1, X2, Y_value

📈 MerQur Output

SPATIAL ERROR MODEL RESULT
─────────────────────────────────────────────

lambda = 0.590 z = 6.38 p < .001 *** Pseudo R^2 = 0.84 (N = 100, k-NN W, k = 5)

💬 Interpretation

This time we modeled spatial dependency in the ERROR term: spatial error autocorrelation is significant (lambda =
0.59, p < .001) — the effect of geographic variables omitted from the model makes neighboring errors correlated.
Unlike SAR, SEM attributes the spread to the error rather than the outcome. In sports science/sports geography, when
the source of spatial autocorrelation is unmeasured geographic factors (infrastructure, access), the correct
specification is SEM; it is chosen by comparison with SAR.

⚠ Common Mistakes

  • λ ≈ 1 → matrix singularity; try a looser W (k=10).
  • The choice of W matrix affects the results — compare KNN and DistanceBand.
  • Projected coordinates (UTM) instead of Lat/Lon may be preferable for accurate distance.

📚 MerQur’a Atıf

Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001

Tüm atıf formatları →

📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.