arXiv:2601.01480stat.MLcs.LG2026-01被引 1

提出可识别数据缺失机制的时序模型,提升交通预测在断连情况下的准确性。

Modeling Information Blackouts in Missing Not-At-Random Time Series Data

  • 构建依赖隐状态的缺失概率模型,区分可忽略与不可忽略缺失
  • 在真实数据上实现4.177 mph的插补误差,优于多数神经网络方法
  • 适合处理有信息性缺失的交通、金融等时间序列场景

交通预测依赖固定传感器网络,常出现连续数据中断。现有方法多将断连视为可忽略缺失,但实际缺失可能与未观测交通状况相关。本文提出一种面向缺失不随机(MNAR)的隐状态空间模型,结合线性交通动态与依赖隐状态的伯努利缺失通道,通过扩展卡尔曼滤波(EKF)与Rauch-Tung-Striebel(RTS)平滑进行推断,参数采用近似期望最大化(EM)学习。在西雅图数据集上,使用300个无泄漏、月度均衡的全时程对齐断连窗口进行评估。MAR-LDS实现4.264 mph的联合插补均方根误差(RMSE),MNAR-LDS进一步降至4.177(改进-0.086);检测器聚类自助法95%置信区间为[-0.182, -0.002]。引入因果一阶预测隐表示后,缺失检测的ROC-AUC从0.685升至0.784。在相同掩码插补协议下,该紧凑概率模型仍优于9个神经架构中的8个,排名第二,仅比最佳结果低1.22%,且无统计显著差异;同时在P95误差、长时段断连误差上更优,存储量仅为神经模型的数百分之一。相较于MAR,MNAR使端到端训练时间翻倍,推理时间增加41%,显现出精度-复杂度-成本权衡。控制状态下依赖状态的断连实验显示,当缺失具有信息量时,性能提升更显著,30分钟预报误差降低6.34%。

原文摘要 · Abstract (English)

Traffic forecasting systems rely on fixed sensor networks that frequently exhibit contiguous blackouts. Such outages are usually treated as ignorable missingness, although dropout can depend on unobserved traffic conditions. We study this possibility with an MNAR-aware latent state-space model that combines linear traffic dynamics with a Bernoulli missingness channel whose probability depends on the latent state. Inference uses an Extended Kalman Filter (EKF) followed by Rauch-Tung-Striebel (RTS) smoothing, and parameters are learned by approximate EM. We evaluate Seattle using a leakage-free, month-balanced set of 300 unique all-horizon-aligned blackout windows. On this benchmark, MAR-LDS attains 4.264 mph pooled imputation RMSE and MNAR-LDS improves it to 4.177 (difference -0.086); the detector-cluster bootstrap 95% interval is [-0.182,-0.002]. A causal one-step predicted latent representation raises missingness ROC-AUC from 0.685 using observed-only features to 0.784. We further test whether this compact probabilistic model remains competitive with substantially larger neural time-series architectures under the identical masked-imputation protocol. MNAR-LDS ranks second in pooled RMSE and outperforms 8 of 9 evaluated neural architectures; it is within 1.22% of the best neural result, with no statistically resolved difference under detector-cluster bootstrap, while achieving lower P95 error, lower long-blackout RMSE, and orders of magnitude fewer stored scalar entries. MNAR roughly doubles end-to-end training time relative to MAR and increases EKF+RTS inference time by 41%, making the accuracy-complexity-cost tradeoff explicit. Controlled state-dependent blackouts further show larger gains when dropout is genuinely informative, including a 6.34% reduction in 30-minute forecast RMSE relative to MAR.

时间序列缺失数据状态空间模型交通预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。