arXiv:2510.02216cs.LGmath.ST2025-10NeurIPS

提出扩散变压器的理论框架,实现时间序列插补的高效与不确定性量化。

Diffusion Transformers for Imputation: Statistical Efficiency and Uncertainty Quantification

  • 基于变换器构建条件得分函数近似理论
  • 导出样本复杂度界并构造缺失值置信区间
  • 揭示缺失模式影响插补效率,适合数据质量差场景

插补方法对提升存在普遍缺失值的实际时间序列数据质量至关重要。近期基于扩散的生成式插补方法在性能上显著优于自回归和传统统计方法。然而,其对缺失值与观测值之间复杂时空依赖关系的捕捉能力仍缺乏理论支撑。本文通过研究条件扩散变压器的统计效率并量化缺失值不确定性,填补该空白。我们基于变换器对条件得分函数的新型近似理论,推导出统计样本复杂度边界,并据此构建缺失值的紧致置信区域。研究发现,插补的效率与精度显著受缺失模式影响。通过模拟验证理论结论,并提出混合掩码训练策略以提升插补性能。

原文摘要 · Abstract (English)

Imputation methods play a critical role in enhancing the quality of practical time-series data, which often suffer from pervasive missing values. Recently, diffusion-based generative imputation methods have demonstrated remarkable success compared to autoregressive and conventional statistical approaches. Despite their empirical success, the theoretical understanding of how well diffusion-based models capture complex spatial and temporal dependencies between the missing values and observed ones remains limited. Our work addresses this gap by investigating the statistical efficiency of conditional diffusion transformers for imputation and quantifying the uncertainty in missing values. Specifically, we derive statistical sample complexity bounds based on a novel approximation theory for conditional score functions using transformers, and, through this, construct tight confidence regions for missing values. Our findings also reveal that the efficiency and accuracy of imputation are significantly influenced by the missing patterns. Furthermore, we validate these theoretical insights through simulation and propose a mixed-masking training strategy to enhance the imputation performance.

时间序列扩散模型插补不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。