让机器学习模型正确使用调查数据设计信息,避免偏差和低估不确定性。
Survey-aware Machine Learning: A Guideline for Valid Population Health Inference based on Scoping Review
- 引入调查设计元数据,分九步整合到机器学习全流程
- 基于16篇文献的综述,验证方法可减少估计偏差并提升公平性评估可靠性
- 适合做人群健康推断的研究者,尤其关注统计有效性和代表性
基于复杂健康调查(如国家健康与营养检查调查,NHANES)训练的机器学习模型常忽略初级抽样单元、分层变量和抽样权重,违背标准评估方法的独立性假设。这导致估计偏差、不确定性被低估,公平性评估无法反映真实人群差异。本文提出调查感知机器学习(SaML),提供九步指南,将调查设计元数据贯穿于机器学习全生命周期。通过16篇方法论文的范围综述,总结了加权模型训练、基于设计的交叉验证及调查调整性能评估的现有工作,并识别出超参数调优与部署环节的空白。提供针对不同分析目标的任务特定指导,帮助研究者实现基于调查数据的有效人口推断。
原文摘要 · Abstract (English)
Machine Learning (ML) models trained on complex health surveys such as the National Health and Nutrition Examination Survey (NHANES) often ignore primary sampling units, stratification variables, and sampling weights. This practice violates the independence assumptions of standard evaluation methods. As a result, estimates become biased, uncertainty is underestimated, and fairness assessments fail to reflect population-level disparities. We propose Survey-aware Machine Learning (SaML), a nine-step guideline that incorporates survey design metadata across the ML lifecycle. Through a scoping review of 16 methodological papers, we summarize existing work on weighted model training, design-based cross-validation, and survey-adjusted performance evaluation. We also identify gaps in hyperparameter tuning and deployment. We provide task-specific guidance that clarifies which steps are required for different analytical objectives. SaML provides a checklist for valid population inference from survey data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。