arXiv:2501.10910cs.LGstat.ML2025-01被引 8

用对比学习与注意力机制,提升非随机缺失数据的填补效果。

DeepIFSAC: Deep Imputation of Missing Values Using Feature and Sample Attention within Contrastive Framework

  • 引入行与列注意力,建模特征与样本间关系以重构缺失值。
  • 在12个数据集上表现优于11种主流方法,高缺失率下仍稳定有效。
  • 适合处理非随机缺失的医疗等真实场景表格数据。

现实世界表格数据中不同模式和比率的缺失值给构建可靠数据驱动模型带来挑战。传统统计与机器学习方法在缺失率高且非随机时效果不佳。本文提出一种新框架,利用行与列注意力建模特征间与样本间关系,实现缺失值重建。通过在对比学习框架中引入CutMix数据增强,提升缺失值估计的不确定性建模能力。在12个多样化表格数据集上,对比11种先进统计、机器学习与深度填补方法,本方法在10%至90%缺失率及三种缺失类型下平均性能排名领先,尤其在非随机缺失情形表现突出。进一步在真实电子健康记录下游患者分类任务中验证,所填补数据质量更高。研究强调表格数据异质性,建议根据缺失类型与数据特性选择适配的填补方法。

原文摘要 · Abstract (English)

Missing values of varying patterns and rates in real-world tabular data pose a significant challenge in developing reliable data-driven models. The most commonly used statistical and machine learning methods for missing value imputation may be ineffective when the missing rate is high and not random. This paper explores row and column attention in tabular data as between-feature and between-sample attention in a novel framework to reconstruct missing values. The proposed method uses CutMix data augmentation within a contrastive learning framework to improve the uncertainty of missing value estimation. The performance and generalizability of trained imputation models are evaluated in set-aside test data folds with missing values. The proposed framework is compared with 11 state-of-the-art statistical, machine learning, and deep imputation methods using 12 diverse tabular data sets. The average performance rank of our proposed method demonstrates its superiority over the state-of-the-art methods for missing rates between 10% and 90% and three missing value types, especially when the missing values are not random. The quality of the imputed data using our proposed method is compared in a downstream patient classification task using real-world electronic health records. This paper highlights the heterogeneity of tabular data sets to recommend imputation methods based on missing value types and data characteristics.

数据填补对比学习医疗数据注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。