arXiv:2503.16644cs.LGcs.HC2025-03被引 2

调研70位工程师发现,多数人处理缺失数据不科学,影响模型可靠性。

To impute or not to impute: How machine learning modelers treat missing data

  • 通过问卷调研70名从业者,分析缺失数据处理方法选择行为
  • 超半数研究者未基于数据特征科学选择处理策略
  • 建议加强教育、标准化报告和开发分析工具

缺失数据在表格型机器学习(ML)中普遍存在,不同处理方法会显著影响模型训练结果。然而,目前尚不清楚机器学习研究人员与工程师如何选择缺失数据处理方法及其影响因素。为此,我们对70名机器学习研究人员和工程师进行了调查。结果显示,大多数参与者在缺失数据处理上并未做出知情决策,这可能严重影响其训练的机器学习模型的有效性。我们呼吁加强缺失数据方面的教育,推动更标准化的缺失数据报告,并开发更好的缺失数据分析工具。

原文摘要 · Abstract (English)

Missing data is prevalent in tabular machine learning (ML) models, and different missing data treatment methods can significantly affect ML model training results. However, little is known about how ML researchers and engineers choose missing data treatment methods and what factors affect their choices. To this end, we conducted a survey of 70 ML researchers and engineers. Our results revealed that most participants were not making informed decisions regarding missing data treatment, which could significantly affect the validity of the ML models trained by these researchers. We advocate for better education on missing data, more standardized missing data reporting, and better missing data analysis tools.

缺失数据机器学习调研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。