提出F3I方法,在缺失数据下保持分布一致性并提升分类性能。
Handling Missing Data in Downstream Tasks With Distribution-Preserving Guarantees
- 基于KNN迭代优化,通过可微凹函数学习邻居权重以保分布
- 在药物重定位和手写数字识别任务中表现优于现有方法
- 理论证明对多种缺失机制均具分布保持性,适合高维数据
缺失特征值是分类等下游机器学习任务的重大挑战。现有插补方法在高维数据上计算耗时,且对数据分布保持和插补质量缺乏理论保障,尤其针对非随机缺失机制。本文提出F3I方法,基于K近邻的迭代改进,通过优化一种新颖的凹可微目标函数来学习邻居特定权重,以保留非缺失值上的数据分布。F3I可与任意分类器架构联合训练。我们进一步对F3I在多种缺失机制下的插补质量和分布保持性提供了理论分析。实验表明,F3I在多个插补与分类任务中表现优异,应用于药物重定位和手写数字识别数据集。
原文摘要 · Abstract (English)
Missing feature values are a significant hurdle for downstream machine-learning tasks such as classification. However, imputation methods for classification might be time-consuming for high-dimensional data, and offer few theoretical guarantees on the preservation of the data distribution and imputation quality, especially for not-missing-at-random mechanisms. First, we propose an imputation approach named F3I based on the iterative improvement of a K-nearest neighbor imputation, where neighbor-specific weights are learned through the optimization of a novel concave, differentiable objective function related to the preservation of the data distribution on non-missing values. F3I can then be chained to and jointly trained with any classifier architecture. Second, we provide a theoretical analysis of imputation quality and data distribution preservation by F3I for several types of missing mechanisms. Finally, we demonstrate the superior performance of F3I on several imputation and classification tasks, with applications to drug repurposing and handwritten-digit recognition data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。