Data Leakage Inflates Generalizability of Power Outage Prediction Models
作者: Yamil Essus, Ranga Raju Vatsavai, Benjamin Rachunok
分类: cs.LG
发布日期: 2026-08-25
💡 一句话要点
揭示数据泄露对电力中断预测模型泛化能力的影响
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 电力中断预测 数据泄露 模型泛化 空间自相关 时间自相关 GeoAI 评估方法 基础设施风险
📋 核心要点
- 现有电力中断预测模型的评估方法未能真实反映其在新条件下的泛化能力,存在数据泄露问题。
- 论文通过比较不同的测试选择策略,提出了更为现实的评估方法,以提高模型的泛化能力。
- 实验结果显示,随机划分的模型性能被高估,空间和时间保持实验中模型表现显著下降,且GeoAI模型嵌入的效果有限。
📝 摘要(中文)
电力中断预测模型在气候驱动的基础设施风险评估中越来越常用,但当前的评估实践掩盖了这些模型在新条件下的泛化能力。本文识别了影响模型泛化能力的三种常见方法选择,并通过对2018至2023年美国东海岸的公开数据进行比较,评估了不同方法决策对预测性能的影响。研究发现,随机训练-测试划分的结果受到空间和时间自相关的影响,模型在空间和时间保持实验下的预测准确性显著下降,且在事件级别的迁移能力较差。这表明,当前的电力中断预测模型在实际应用中价值有限,未来需要改进数据覆盖和评估协议。
🔬 方法详解
问题定义:本文旨在解决电力中断预测模型在新条件下的泛化能力不足的问题,现有方法在评估时存在数据泄露现象,导致模型性能被高估。
核心思路:通过识别并比较三种常见的评估方法,论文提出了更为严谨的测试选择策略,以更真实地反映模型在实际应用中的表现。
技术框架:研究采用了来自2018至2023年美国东海岸的公开数据,结合天气重分析和土地覆盖数据,以及GeoAI基础模型的嵌入特征,构建了多个模型进行比较。
关键创新:论文的创新在于通过空间和时间保持实验揭示了模型性能的真实情况,强调了现有评估方法的局限性,并指出了数据覆盖和评估协议改进的必要性。
关键设计:研究中采用了多种测试选择策略,包括随机划分、留一州和留一事件的设计,评估模型在不同条件下的表现,特别关注空间和事件级别的迁移能力。
🖼️ 关键图片
📊 实验亮点
实验结果表明,随机训练-测试划分的模型性能被高估,空间和时间保持实验中模型的预测准确性显著下降,常常无法超越简单的基线模型。GeoAI模型嵌入的改进效果有限,主要体现在空间泛化上,事件级别的迁移能力仍然较差。
🎯 应用场景
该研究的潜在应用领域包括电力基础设施的风险评估和气候变化影响分析。通过改进电力中断预测模型的评估方法,可以为决策者提供更可靠的工具,以应对气候变化带来的挑战,提升基础设施的韧性和安全性。
📄 摘要(原文)
Power outage prediction models are increasingly used in assessments of climate-driven infrastructure risk, yet current evaluation practices obscure whether these models generalize to the novel conditions such applications require. We identify three common methodological choices in power outage prediction models that influence their ability to generalize across spatial, temporal, and event-based settings. We compare the predictive performance impacts of different methodological decisions using publicly available data for the U.S. East Coast from 2018 to 2023 and feature sets derived from weather reanalysis and land-cover data, and embeddings from a GeoAI foundation model (Prithvi WxC). Specifically, we assess model performance under multiple test selection strategies, including unfiltered random splits, leave-one-state-out, and leave-one-event-out designs, which increasingly approximate real-world deployment conditions. While random train-test splits yield strong performance, we show that these results are inflated by spatial and temporal autocorrelation. Under spatial and temporal holdout experiments, predictive accuracy degrades substantially, with models often failing to outperform a simple null baseline. Incorporating GeoAI foundation model embeddings yields limited and inconsistent improvements, primarily for spatial generalization, and does not resolve poor event-level transferability. These findings suggest that, given current data availability and evaluation practices, publicly trained outage prediction models offer limited and uncertain operational value. Progress will likely require improved data coverage, more realistic evaluation protocols, and a shift in focus from marginal modeling advances toward addressing structural data constraints.