arXiv:2504.04243cs.LGcs.AI2025-04被引 4

研究发现,医疗AI标签不确定性会导致预测结果差异巨大。

Perils of Label Indeterminacy: A Case Study on Prediction of Neurological Recovery After Cardiac Arrest

  • 提出标签不确定性的概念,揭示医疗AI评估中的隐性假设风险
  • 实证显示:标签未知患者预测结果差异可达显著水平
  • 提醒临床AI开发者关注评估偏差,尤其在生死决策场景

设计用于辅助人类决策的AI系统通常需要标注数据来训练和评估监督模型。然而,这些标签往往不可得,而不同估算方法依赖无法验证的假设或任意选择。本文引入标签不确定性的概念,并探讨其在高风险AI辅助决策中的重要影响。我们以心脏骤停后昏迷患者复苏恢复预测为案例进行实证研究。结果显示,尽管在标签已知患者上模型表现相似,但在标签未知患者上的预测却存在巨大差异。该现象揭示了标签不确定性的关键伦理问题。最后,我们讨论了在评估、报告与设计中应采取的改进措施。

原文摘要 · Abstract (English)

The design of AI systems to assist human decision-making typically requires the availability of labels to train and evaluate supervised models. Frequently, however, these labels are unknown, and different ways of estimating them involve unverifiable assumptions or arbitrary choices. In this work, we introduce the concept of label indeterminacy and derive important implications in high-stakes AI-assisted decision-making. We present an empirical study in a healthcare context, focusing specifically on predicting the recovery of comatose patients after resuscitation from cardiac arrest. Our study shows that label indeterminacy can result in models that perform similarly when evaluated on patients with known labels, but vary drastically in their predictions for patients where labels are unknown. After demonstrating crucial ethical implications of label indeterminacy in this high-stakes context, we discuss takeaways for evaluation, reporting, and design.

医疗AI标签不确定性高风险决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。