arXiv:2511.01196stat.MLcs.AI2025-11综述被引 3

系统梳理缺失数据填补方法,连接统计学与现代机器学习。

An Interdisciplinary and Cross-Task Review on Missing Data Imputation

  • 按缺失机制和任务类型分类,整合经典与深度学习方法
  • 涵盖张量、时间序列等复杂数据类型填补技术
  • 适合跨领域研究者参考,尤其关注下游任务融合

缺失数据是数据科学中的基础挑战,严重阻碍医疗、生物信息学、社会科学、电子商务和工业监控等领域的分析与决策。尽管历经数十年研究并涌现出众多填补方法,但各领域文献仍分散割裂,亟需全面整合以连接统计基础与现代机器学习进展。本文系统回顾核心概念——缺失机制、单次与多次填补、不同填补目标——并分析跨领域问题特征。全面分类填补方法,涵盖经典技术(如回归、EM算法)及现代方法(低秩/高秩矩阵补全、自编码器、GANs、扩散模型、图神经网络)和大语言模型。特别关注张量、时间序列、流数据、图结构数据、类别数据及多模态数据等复杂数据类型的填补方法。除方法外,还探讨填补与分类、聚类、异常检测等下游任务的集成,分析顺序流程与联合优化框架。评估理论保证、基准资源与评估指标,并指出关键挑战与未来方向,强调模型选择与超参数优化、联邦学习支持的隐私保护填补,以及跨领域、跨数据类型可迁移模型的构建,为未来研究提供路线图。

原文摘要 · Abstract (English)

Missing data is a fundamental challenge in data science, significantly hindering analysis and decision-making across a wide range of disciplines, including healthcare, bioinformatics, social science, e-commerce, and industrial monitoring. Despite decades of research and numerous imputation methods, the literature remains fragmented across fields, creating a critical need for a comprehensive synthesis that connects statistical foundations with modern machine learning advances. This work systematically reviews core concepts-including missingness mechanisms, single versus multiple imputation, and different imputation goals-and examines problem characteristics across various domains. It provides a thorough categorization of imputation methods, spanning classical techniques (e.g., regression, the EM algorithm) to modern approaches like low-rank and high-rank matrix completion, deep learning models (autoencoders, GANs, diffusion models, graph neural networks), and large language models. Special attention is given to methods for complex data types, such as tensors, time series, streaming data, graph-structured data, categorical data, and multimodal data. Beyond methodology, we investigate the crucial integration of imputation with downstream tasks like classification, clustering, and anomaly detection, examining both sequential pipelines and joint optimization frameworks. The review also assesses theoretical guarantees, benchmarking resources, and evaluation metrics. Finally, we identify critical challenges and future directions, emphasizing model selection and hyperparameter optimization, the growing importance of privacy-preserving imputation via federated learning, and the pursuit of generalizable models that can adapt across domains and data types, thereby outlining a roadmap for future research.

缺失数据综述机器学习数据修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。