arXiv:2605.06355cs.LGstat.ML2026-05

让自回归模型直接处理缺失数据,提升不完整数据下的生成与预测能力。

Order-Agnostic Autoregressive Modelling with Missing Data

论文配图:Order-Agnostic Autoregressive Modelling with Missing Data
图 1 · 摘自论文原文
  • 将缺失数据视为随机缺失,重构训练过程以支持不完整数据集
  • 在真实数据集上优于主流插补方法,尤其在高缺失率下表现更优
  • 可主动选择最有效的缺失变量进行补全,适合数据采集优化场景

顺序无关的自回归模型在深度生成建模中表现出色,但在不完整数据场景下的应用仍较少。本文从缺失数据视角重新审视该类模型:首先表明,其在完整数据上的标准训练隐式地执行了随机缺失机制下的插补,从而在高缺失率情况下具备强健的外样本插补性能;其次提出首个在一般缺失机制下直接于不完整数据集上训练的合理框架;第三,利用其摊销条件密度估计能力,实现主动信息获取——即按序选择对下游预测或推断最有价值的缺失变量。在一系列真实世界基准测试中,我们的缺失感知顺序无关自回归模型(MO-ARM)持续优于现有插补基线。

原文摘要 · Abstract (English)

Order-Agnostic autoregressive models have demonstrated strong performance in deep generative modeling, yet their use in settings with incomplete data remains largely unexplored. In this work, we reinterpret them through the lens of missing data. First, we show that their standard training procedure on fully observed data implicitly performs imputation under a missing completely at random mechanism, resulting in robust out-of-sample imputation performance in settings with high missingness. Second, we introduce the first principled framework for training them directly on incomplete datasets under general missingness mechanisms. Third, we leverage their amortized conditional density estimation to perform active information acquisition, i.e., sequentially selecting the most informative missing variables for downstream prediction or inference. Across a suite of real-world benchmarks, our Missingness-Aware Order-Agnostic Autoregressive Model (MO-ARM) consistently outperforms established imputation baselines.

生成模型缺失数据主动学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。