SAR (Spatial Autoregressive Lag)
v1.0.2
Health Sciences context — Spatial Lag Model. A regression that accounts for the spatial neighborhood structure of diyabet_prevalans_pct. Spatial weight matrix: K-nearest neighbor (k=5), row-standardized.
🎯 What is it for?
Classic OLS assumes the observations are independent — but in lat/lon data, nearby points take on similar values (spatial autocorrelation). SAR (Spatial Autoregressive Lag) incorporates this structure into the model; otherwise the standard errors of the OLS β estimates come out too small (type-I error).
📌 When is it used?
- Spatial data (with lat/lon) + a continuous DV in the Health Sciences
- Moran’s I p < .05 — evidence of spatial autocorrelation
- Spatial clustering in OLS residuals (hotspot over DBSCAN)
⚙ Assumptions
- Lat/lon coordinates ({lat, lon}).
- Continuous DV (diyabet_prevalans_pct).
- Spatial weight matrix design (KNN k=5 — default, with dist.band as an alternative).
- The ρ parameter is stable within 0-1; there is a risk of fragility at the boundary.
📊 How to Run It in MerQur
Panel assignments (form fields in the program):
- Columns:
{'y': 'vaka_orani', 'x': ['sosyoekonomik', 'yas_ortalama'], 'lat': 'lat', 'lon': 'lon'} - Parameters:
{'weights': 'knn', 'k': 5}
📊 Sample Dataset — Health Sciences
ℹ Note: The scenario, MerQur output and interpretation below were produced by actually running the real example dataset in MerQur. Numeric results on your own data will differ; the goal is to show how the analysis is set up and interpreted end-to-end.
🎬 Example File
This analysis is demonstrated on the following example dataset for Health Sciences:
Tip/105_spatial_sar_COVID_case_ratio.xlsx
🎬 Scenario
When modeling province-level case ratio we handle spatial spillover with SAR.
For neighborhood/contagion effects, SAR is appropriate.
⚙️ Variable Selection
- Dependent variable: case_ratio
- Predictor(s): socioeconomic
- Predictor(s): age_mean
- Latitude: lat
- Longitude: lon
Data Preview (First 5 Rows)
| province_id | lat | lon | socioeconomic | age_mean | case_ratio |
|---|---|---|---|---|---|
| 1.0 | 41.378 | 40.0929 | 33.5 | 34.0 | 510.9 |
| 2.0 | 37.9454 | 38.0209 | 59.6 | 42.9 | 638.7 |
| 3.0 | 40.0527 | 27.3987 | 23.5 | 44.2 | 530.3 |
| 4.0 | 37.8009 | 33.2798 | 73.1 | 50.3 | 537.9 |
| 5.0 | 38.4268 | 31.0605 | 56.5 | 36.5 | 520.4 |
n = 100 · Columns: province_id, lat, lon, socioeconomic, age_mean, case_ratio
📈 MerQur Output
─────────────────────────────────────────────
rho = 0.319 z = 3.95 p < .001 *** Pseudo R^2 = 0.77 (N = 100, k-NN W, k = 5)
outcome: case_ratio, predictors: socioeconomic, age_mean
💬 Interpretation
When modeling the province-level case ratio (case_ratio), we handled spatial spillover (the effect of neighboring
provinces) with SAR: the spatial lag parameter is significant and positive (rho = 0.32, p < .001) — a province’s
case ratio is related to its neighbors’, a “cluster/spread” pattern. SAR incorporates spatial dependency into the
model; if ignored, standard errors are biased. In medicine/epidemiology it is the right method for modeling the
geographic spread of disease-rate indicators (neighborhood/contagion effect).
⚠ Common Mistakes
- ρ ≈ 1 → matrix singularity; try a looser W (k=10).
- The choice of W matrix affects the results — compare KNN and DistanceBand.
- Projected coordinates (UTM) instead of Lat/Lon may be preferable for accurate distance.
📚 MerQur’a Atıf
Örücü, Ö. K. (2026). MerQur: Integrated Academic Data Analysis & Reporting Platform [Computer software] (Version 1.0.0). https://doi.org/10.53463/merqur.2026001
📝 Üretim Notu — Bu sayfadaki örnek veri sentetik olarak üretilmiştir (sabit SEED=42, generator: samples/Ileri_Duzey_v102/_generate_v102_datasets.py). Sayfa içeriği Anthropic Claude desteği ile hazırlanmış, akademik doğruluk yazar tarafından kontrol edilmiştir.