arXiv:2506.09258cs.LGstat.ML2025-06被引 4

CFMI用流模型解决缺失数据填补,效果优于传统与深度学习方法。

CFMI: Flow Matching for Missing Data Imputation

  • 结合连续归一化流与条件建模,解决多重填补的复杂性问题。
  • 在24个数据集上表现超越9种经典与前沿方法,涵盖多种指标。
  • 计算高效,适合低维到高维数据,零样本时间序列填补也出色。

我们提出条件流匹配填补(CFMI),一种通用的缺失数据填补新方法。该方法融合连续归一化流、流匹配与共享条件建模,克服传统多重填补的不可计算性。在24个中小型表格数据集上,与九种经典及前沿方法对比显示,CFMI在多种指标上达到或超过现有技术表现。应用于零样本时间序列填补时,其精度媲美基于扩散模型的方法,但计算效率更高。总体而言,CFMI在低维数据上不弱于传统方法,且可扩展至高维场景,性能与深度学习方法相当甚至更优,适用于多种数据类型与维度,是广泛适用的填补首选方案。

原文摘要 · Abstract (English)

We introduce conditional flow matching for imputation (CFMI), a new general-purpose method to impute missing data. The method combines continuous normalising flows, flow-matching, and shared conditional modelling to deal with intractabilities of traditional multiple imputation. Our comparison with nine classical and state-of-the-art imputation methods on 24 small to moderate-dimensional tabular data sets shows that CFMI matches or outperforms both traditional and modern techniques across a wide range of metrics. Applying the method to zero-shot imputation of time-series data, we find that it matches the accuracy of a related diffusion-based method while outperforming it in terms of computational efficiency. Overall, CFMI performs at least as well as traditional methods on lower-dimensional data while remaining scalable to high-dimensional settings, matching or exceeding the performance of other deep learning-based approaches, making it a go-to imputation method for a wide range of data types and dimensionalities.

数据填补流模型时间序列高维数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。