arXiv:2506.02306cs.LGstat.ML2025-06ICML被引 6

通过缺失模式与上下文信息提升表格数据填补效果

CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data Imputation

  • 采用中位数截断复制掩码策略捕捉缺失规律
  • 在非随机缺失下提升7.8%的R²表现
  • 适合处理带语义信息的复杂表格数据

我们提出CACTI,一种基于掩码自编码的表格数据填补方法,利用缺失模式结构和上下文信息。该方法采用新颖的中位数截断复制掩码训练策略,使模型学习真实缺失模式,同时结合列名和文本描述捕捉特征间的语义关系,以更好表示特征依赖。双重归纳偏置使CACTI在多种数据集和缺失条件下均优于现有最优方法:在缺失非随机、随机和完全随机情况下,平均R²分别提升13.4%、6.1%和5.3%,整体超越次优方法7.8%。结果表明,利用特定数据集的上下文信息和缺失模式可显著提升填补性能。

原文摘要 · Abstract (English)

We present CACTI, a masked autoencoding approach for imputing tabular data that leverages the structure in missingness patterns and contextual information. Our approach employs a novel median truncated copy masking training strategy that encourages the model to learn from empirical patterns of missingness while incorporating semantic relationships between features - captured by column names and text descriptions - to better represent feature dependence. These dual sources of inductive bias enable CACTI to outperform state-of-the-art methods - an average $R^2$ gain of 7.8% over the next best method (13.4%, 6.1%, and 5.3% under missing not at random, at random and completely at random, respectively) - across a diverse range of datasets and missingness conditions. Our results highlight the value of leveraging dataset-specific contextual information and missingness patterns to enhance imputation performance.

数据填补表格数据缺失模式自编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。