SAR (Spatial Autoregressive Lag)
v1.0.2
Education Sciences context — Spatial Lag Model. A regression that accounts for the spatial neighborhood structure of okul_basari_skoru. Spatial weight matrix: K-nearest neighbor (k=5), row standardized.
🎯 What is it for?
Classic OLS assumes the observations are independent — but in lat/lon data, nearby points take on similar values (spatial autocorrelation). SAR (Spatial Autoregressive Lag) incorporates this structure into the model; otherwise the standard errors of the OLS β estimates come out too small (type-I error).
📌 When is it used?
- Spatial data (with lat/lon) + a continuous DV in Education Sciences
- Moran’s I p < .05 — evidence of spatial autocorrelation
- Spatial clustering in the OLS residuals (hotspot over DBSCAN)
⚙ Assumptions
- Lat/lon coordinates ({lat, lon}).
- Continuous DV (school_success_score).
- Spatial weight matrix design (KNN k=5 — default, with dist.band as an alternative).
- The ρ parameter is stable between 0 and 1; there is a risk of fragility near the boundary.
📊 How to Run in MerQur
Panel assignments (form fields in the program):
- Columns:
{'y': 'birim_id', 'x': ['X1', 'X2'], 'lat': 'lat', 'lon': 'lon'} - Parameters:
{'weights': 'knn', 'k': 5}
📊 Sample Dataset — Education Sciences
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Education Sciences:
Egitim_Bilimleri/105_spatial_sar_spatial.xlsx
🎬 Scenario
Suppose we mapped an educational indicator across 100 geographic units such
as school catchment areas and noticed that neighboring areas tend to have
similar values. We have the coordinates lat and lon, two area-level
predictors X1 and X2, and the outcome Y_value. A Spatial Lag (SAR) model is
appropriate because it adds a spatially lagged term for the outcome,
capturing the idea that an area’s result is influenced by the results of its
neighbors, which would bias an ordinary regression. This lets us estimate
the predictor effects while explicitly accounting for spatial spillover
between schools.
⚙️ Variable Selection
- Coordinates: lat, lon
- Predictors: X1, X2
- Dependent variable: Y_value
Data Preview (First 5 Rows)
| unit_id | lat | lon | X1 | X2 | Y_value |
|---|---|---|---|---|---|
| 1.0 | 40.729 | 36.2513 | 58.0 | 18.3 | 375.9 |
| 2.0 | 37.3016 | 39.779 | 60.9 | 46.8 | 464.0 |
| 3.0 | 38.3563 | 33.0728 | 73.2 | 48.6 | 530.5 |
| 4.0 | 41.2077 | 28.8098 | 82.7 | 15.0 | 197.6 |
| 5.0 | 40.7324 | 29.6924 | 48.9 | 32.5 | 385.3 |
n = 100 · Columns: unit_id, lat, lon, X1, X2, Y_value
📈 MerQur Output
─────────────────────────────────────────────
rho (spatial lag) = 0.34 z = 4.94 p < .001 Pseudo R2 = 0.87
Y ~ X1 + X2 + neighbour Y
💬 Interpretation
Modelling a spatial outcome (e.g. regional school achievement), we accounted for spatial dependence — the
tendency of nearby units to be similar. The SAR (spatial autoregressive) model links a unit’s outcome to its
neighbours’ outcome too (the rho parameter). Here rho = 0.34, significantly positive (z = 4.94, p < .001): the
neighbour effect is real and moderate — a region’s achievement is influenced by neighbouring regions. The model
explains 87% of the variance. Using a spatial model is critical, because ordinary regression gives spurious
significant results when spatial autocorrelation is present. SAR is the right method for spatial educational data
with diffusion and neighbour effects.
⚠ Common Mistakes
- ρ ≈ 1 → matrix singularity; try a looser W (k=10).
- The choice of W matrix affects the results — compare KNN and DistanceBand.
- Projected coordinates (UTM) instead of Lat/Lon may be preferable for accurate distance.
📚 MerQur’a Atıf
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.