针对真实数据中缺失机制复杂问题,提出更鲁棒的处理方法。
Developing robust methods to handle missing data in real-world applications effectively
- 区分MCAR、MAR、MNAR三种缺失机制,设计适应性处理策略。
- 突破传统仅假设MCAR的局限,提升多种场景下数据补全效果。
- 适合医疗、金融等缺失数据普遍的真实应用领域研究者。
缺失数据是表格、传感器、时间序列、图像等多种数据类型中普遍存在的挑战,其成因多样,导致不同的缺失机制。以往研究多集中于缺失完全随机(MCAR)的假设,但缺失随机(MAR)和非随机缺失(MNAR)机制同样常见,却常被忽视。本博士研究提出一个全面的研究框架,深入探讨不同缺失机制的影响,旨在开发能有效应对MCAR、MAR和MNAR特性的稳健方法。通过填补这些空白,该研究增强了对跨行业、多模态数据中缺失数据挑战的理解,并为研究人员与实践者提供切实可行的解决方案,使他们能够更有信心地利用不完整数据集。
原文摘要 · Abstract (English)
Missing data is a pervasive challenge spanning diverse data types, including tabular, sensor data, time-series, images and so on. Its origins are multifaceted, resulting in various missing mechanisms. Prior research in this field has predominantly revolved around the assumption of the Missing Completely At Random (MCAR) mechanism. However, Missing At Random (MAR) and Missing Not At Random (MNAR) mechanisms, though equally prevalent, have often remained underexplored despite their significant influence. This PhD project presents a comprehensive research agenda designed to investigate the implications of diverse missing data mechanisms. The principal aim is to devise robust methodologies capable of effectively handling missing data while accommodating the unique characteristics of MCAR, MAR, and MNAR mechanisms. By addressing these gaps, this research contributes to an enriched understanding of the challenges posed by missing data across various industries and data modalities. It seeks to provide practical solutions that enable the effective management of missing data, empowering researchers and practitioners to leverage incomplete datasets confidently.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。