arXiv:2601.05017cs.LGcs.AI2026-01

统一处理异构特征缺失值,提升表格数据补全精度

HMVI: Unifying Heterogeneous Attributes with Natural Neighbors for Missing Value Inference

  • 基于自然邻居构建统一框架,建模数值与类别特征间依赖
  • 在多个真实数据集上优于现有方法,显著提升下游任务性能
  • 适合存在异构特征缺失的工业级数据补全场景

缺失值填补是机器智能中的基础挑战,高度依赖数据完整性。现有方法通常独立处理数值型与类别型特征,忽视了异构特征间的关键关联。为此,我们提出一种新填补方法,在统一框架中显式建模跨类型特征依赖关系。该方法利用完整与不完整样本,确保表格数据填补的准确性与一致性。大量实验表明,所提方法在性能上超越现有技术,并显著提升下游机器学习任务表现,为现实系统中的缺失数据问题提供稳健解决方案。

原文摘要 · Abstract (English)

Missing value imputation is a fundamental challenge in machine intelligence, heavily dependent on data completeness. Current imputation methods often handle numerical and categorical attributes independently, overlooking critical interdependencies among heterogeneous features. To address these limitations, we propose a novel imputation approach that explicitly models cross-type feature dependencies within a unified framework. Our method leverages both complete and incomplete instances to ensure accurate and consistent imputation in tabular data. Extensive experimental results demonstrate that the proposed approach achieves superior performance over existing techniques and significantly enhances downstream machine learning tasks, providing a robust solution for real-world systems with missing data.

缺失值填补表格数据特征依赖统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。