arXiv:2508.04740cs.LGcs.MS2025-08被引 3

一站式工具,模拟和分析数据缺失机制。

MissMecha: An All-in-One Python Package for Studying Missing Data Mechanisms

  • 支持数值与类别变量的缺失机制仿真
  • 覆盖MCAR、MAR、MNAR三类缺失场景
  • 适合数据质量研究与教学使用

不完整数据是真实世界数据集中的常见挑战,常由复杂且不可观测的缺失机制驱动。模拟缺失已成为理解其对学习与分析影响的标准方法。然而,现有工具分散、机制有限,通常仅关注数值变量,忽视了真实表格数据的异质性。我们提出 MissMecha,一个开源 Python 工具包,用于在 MCAR、MAR 和 MNAR 假设下模拟、可视化和评估缺失数据。它同时支持数值与类别特征,可在混合类型表格数据中实现机制感知的研究。工具包含可视化诊断、MCAR 检测功能以及类型感知的插补评估指标。专为数据质量研究、基准测试和教学设计,提供统一平台,助力研究人员与从业者应对不完整数据问题。

原文摘要 · Abstract (English)

Incomplete data is a persistent challenge in real-world datasets, often governed by complex and unobservable missing mechanisms. Simulating missingness has become a standard approach for understanding its impact on learning and analysis. However, existing tools are fragmented, mechanism-limited, and typically focus only on numerical variables, overlooking the heterogeneous nature of real-world tabular data. We present MissMecha, an open-source Python toolkit for simulating, visualizing, and evaluating missing data under MCAR, MAR, and MNAR assumptions. MissMecha supports both numerical and categorical features, enabling mechanism-aware studies across mixed-type tabular datasets. It includes visual diagnostics, MCAR testing utilities, and type-aware imputation evaluation metrics. Designed to support data quality research, benchmarking, and education,MissMecha offers a unified platform for researchers and practitioners working with incomplete data.

数据缺失Python工具数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。