arXiv:2508.07617cs.HCcs.AI2025-08被引 2

AI拒绝对诊断有帮助,但会漏诊更多病。

Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives

  • AI只在有信心时才给出建议,避免错误影响医生判断。
  • 用拒绝预测后医生整体准确率回升至64%,但漏诊率升18%。
  • 适合关注人机协作安全性的医疗AI研发与临床部署者。

AI有潜力辅助人类决策,但高性能模型在部署时仍可能产生错误预测。这些错误与自动化偏倚(人类过度依赖AI)结合,反而导致更差决策。选择性预测(仅当模型置信度高时输出结果)被提出作为解决方案,其假设是:当AI主动回避时,人类会像无AI介入时一样决策。为验证此假设,我们在临床场景下对259名医生进行用户研究,让他们诊断住院患者。比较了无AI、有AI但无选择性预测、有选择性预测三种情况下的诊断与治疗准确率。结果显示,选择性预测能缓解低质量AI带来的负面影响:相比无AI时的66%准确率(95% CI: 56%-75%),使用错误AI预测时准确率降至56%(95% CI: 46%-66%),而采用选择性预测后恢复至64%(95% CI: 54%-73%)。然而,尽管整体准确率基本保持,选择性预测改变了错误模式——当被告知AI拒绝预测时,医生漏诊率上升18%,漏治率上升35%,比完全无AI输入时更严重。研究强调必须实证检验人类在人机系统中如何与AI互动。

原文摘要 · Abstract (English)

AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deployed. These inaccuracies, combined with automation bias, where humans overrely on AI predictions, can result in worse decisions. Selective prediction, in which potentially unreliable model predictions are hidden from users, has been proposed as a solution. This approach assumes that when AI abstains and informs the user so, humans make decisions as they would without AI involvement. To test this assumption, we study the effects of selective prediction on human decisions in a clinical context. We conducted a user study of 259 clinicians tasked with diagnosing and treating hospitalized patients. We compared their baseline performance without any AI involvement to their AI-assisted accuracy with and without selective prediction. Our findings indicate that selective prediction mitigates the negative effects of inaccurate AI in terms of decision accuracy. Compared to no AI assistance, clinician accuracy declined when shown inaccurate AI predictions (66% [95% CI: 56%-75%] vs. 56% [95% CI: 46%-66%]), but recovered under selective prediction (64% [95% CI: 54%-73%]). However, while selective prediction nearly maintains overall accuracy, our results suggest that it alters patterns of mistakes: when informed the AI abstains, clinicians underdiagnose (18% increase in missed diagnoses) and undertreat (35% increase in missed treatments) compared to no AI input at all. Our findings underscore the importance of empirically validating assumptions about how humans engage with AI within human-AI systems.

人机协作医疗AI自动化偏倚选择性预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。