ZIGAT-MAE: A Context-Conditioned Zero-Inflated Graph Attention Network with Masked Autoencoders for Spatio-Temporal Urban Crime-Risk Forecasting
Resumen
Predicting urban crime struggles with sparse data, spatial dependencies, and variable log reliability. We propose ZIGAT-MAE, a graph attention network on H3 hexagonal grids. It fuses dynamic, static, and environmental data via dual-layer GAT and FiLM conditioning, outputting Zero-Inflated Poisson (ZIP) parameters through dual heads to handle differing data streams, optimized by a sparsity-conscious loss. The model uses multi-city masked autoencoder (MAE) pretraining to reconstruct static features and "ghost-imputation" to drop untrustworthy data. Evaluated across 9 Mexican metros, ZIGAT-MAE achieves a 0.942 ROC-AUC, 194x baseline average precision, and excellent calibration errors (under 10^-3). Targeting the highest-risk 5% of areas captures 72% of incidents, beating historical baselines. Ablations confirm MAE pretraining and ZIP fine-tuning are synergistic. Extreme crime clustering (Gini 0.96; top 5% of areas see 89% of crime) validates the architecture for strategic resource deployment.