arXiv:2608.24665cs.LG2026-08

现有停电预测模型泛化能力被数据泄露夸大,真实场景下效果有限。

Data Leakage Inflates Generalizability of Power Outage Prediction Models

论文配图:Data Leakage Inflates Generalizability of Power Outage Prediction Models
图 1 · 摘自论文原文
  • 对比多种测试策略,发现随机划分导致性能虚高
  • 空间时间留一法下准确率大幅下降,常不如基线模型
  • 地理大模型嵌入提升有限,难以解决事件级迁移问题

停电预测模型在评估气候驱动的基础设施风险中日益重要,但当前评估方法掩盖了其在新条件下的泛化能力。本文识别出三种影响模型跨空间、时间与事件泛化能力的方法论选择。基于2018至2023年美国东海岸公开数据,结合气象再分析、土地覆盖特征及地理人工智能基础模型(Prithvi WxC)的嵌入表示,比较不同方法学决策对预测性能的影响。评估涵盖无筛选随机划分、留一州、留一事件等测试策略,逐步逼近实际部署场景。尽管随机划分下表现优异,但结果受空间与时间自相关性扭曲。在空间和时间留出实验中,预测准确率显著下降,多数模型甚至无法超越简单零假设基线。引入地理大模型嵌入仅带来有限且不一致的改进,主要体现在空间泛化,未能解决事件级别迁移问题。这表明,受限于当前数据覆盖与评估实践,公开训练的停电预测模型操作价值有限且不确定。进步需依赖更完善的数据、更真实的评估协议,并从边际建模改进转向结构性数据约束的根本解决。

原文摘要 · Abstract (English)

Power outage prediction models are increasingly used in assessments of climate-driven infrastructure risk, yet current evaluation practices obscure whether these models generalize to the novel conditions such applications require. We identify three common methodological choices in power outage prediction models that influence their ability to generalize across spatial, temporal, and event-based settings. We compare the predictive performance impacts of different methodological decisions using publicly available data for the U.S. East Coast from 2018 to 2023 and feature sets derived from weather reanalysis and land-cover data, and embeddings from a GeoAI foundation model (Prithvi WxC). Specifically, we assess model performance under multiple test selection strategies, including unfiltered random splits, leave-one-state-out, and leave-one-event-out designs, which increasingly approximate real-world deployment conditions. While random train-test splits yield strong performance, we show that these results are inflated by spatial and temporal autocorrelation. Under spatial and temporal holdout experiments, predictive accuracy degrades substantially, with models often failing to outperform a simple null baseline. Incorporating GeoAI foundation model embeddings yields limited and inconsistent improvements, primarily for spatial generalization, and does not resolve poor event-level transferability. These findings suggest that, given current data availability and evaluation practices, publicly trained outage prediction models offer limited and uncertain operational value. Progress will likely require improved data coverage, more realistic evaluation protocols, and a shift in focus from marginal modeling advances toward addressing structural data constraints.

停电预测泛化能力数据泄露地理AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。