提出分层压力测试,揭示时间序列填补中的平稳性偏差。
The Stationarity Bias: Stratified Stress-Testing for Time-Series Imputation in Regulated Dynamical Systems
- 按平稳与波动阶段分层评估填补效果,避免误判模型性能。
- 线性填补在平稳期表现最优,但波动期形状失真严重。
- 深度模型更适配医疗等安全关键场景的动态填补需求。
时间序列填补基准普遍采用均匀随机掩码和无形状感知指标(如MSE、RMSE),隐含地按状态出现频率加权评估。在具有主导吸引子的系统中——如稳态生理、正常工业运行、稳定网络流量——这导致系统性偏差:简单方法因在低熵易处理状态下表现优异而显得优越。本文正式定义此为‘平稳性偏差’,并提出‘分层压力测试’,将评估划分为平稳与瞬变阶段。以连续血糖监测(CGM)为实验平台,因其有严格真实强迫函数(进餐、胰岛素)可精确定位状态,得出三项结论:(i) 平稳效率:线性插值在稳定区间达当前最优重建效果,证实复杂架构在低熵环境下冗余;(ii) 瞬变保真度:在关键瞬变期(餐后峰值、低血糖事件),线性方法形态保真度显著下降(DTW),远超其RMSE所反映,称之为‘RMSE幻象’,即点误差低却破坏信号形状;(iii) 阶段条件模型选择:深度学习模型在瞬变期同时保持点精度与形态完整性,对安全关键下游任务至关重要。进一步基于临床试验数据提取缺失分布,并施加于完整训练数据,防止模型依赖不现实的干净观测,提升真实缺失下的鲁棒性。该框架适用于任何以常规平稳为主、关键瞬变为辅的受控系统。
原文摘要 · Abstract (English)
Time-series imputation benchmarks employ uniform random masking and shape-agnostic metrics (MSE, RMSE), implicitly weighting evaluation by regime prevalence. In systems with a dominant attractor -- homeostatic physiology, nominal industrial operation, stable network traffic -- this creates a systematic \emph{Stationarity Bias}: simple methods appear superior because the benchmark predominantly samples the easy, low-entropy regime where they trivially succeed. We formalize this bias and propose a \emph{Stratified Stress-Test} that partitions evaluation into Stationary and Transient regimes. Using Continuous Glucose Monitoring (CGM) as a testbed -- chosen for its rigorous ground-truth forcing functions (meals, insulin) that enable precise regime identification -- we establish three findings with broad implications:(i)~Stationary Efficiency: Linear interpolation achieves state-of-the-art reconstruction during stable intervals, confirming that complex architectures are computationally wasteful in low-entropy regimes.(ii)~Transient Fidelity: During critical transients (post-prandial peaks, hypoglycemic events), linear methods exhibit drastically degraded morphological fidelity (DTW), disproportionate to their RMSE -- a phenomenon we term the \emph{RMSE Mirage}, where low pointwise error masks the destruction of signal shape.(iii)~Regime-Conditional Model Selection: Deep learning models preserve both pointwise accuracy and morphological integrity during transients, making them essential for safety-critical downstream tasks. We further derive empirical missingness distributions from clinical trials and impose them on complete training data, preventing models from exploiting unrealistically clean observations and encouraging robustness under real-world missingness. This framework generalizes to any regulated system where routine stationarity dominates critical transients.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。