arXiv:2502.03482eess.IVcs.AI2025-02被引 3

研究医生如何合理依赖AI诊断前列腺癌MRI,发现反馈不能提升效果但提前看AI结果会更信任。

Can Domain Experts Rely on AI Appropriately? A Case Study on AI-Assisted Prostate Cancer MRI Diagnosis

  • 设计两种临床流程:先独立诊断再看AI,或先看性能反馈再参考AI。
  • 人机团队整体优于人单独判断,但仍弱于纯AI,因医生未充分信任AI。
  • 展示历史表现数据没提升团队效果,但提前见AI结果能促使医生跟从。

尽管人机决策协作日益受关注,但与领域专家(如放射科医生)开展实验仍罕见,主要因合作复杂且难以构建真实场景。本文深入与前列腺癌MRI诊断的放射科医生合作,基于现有教学工具开发界面,并开展两项实验,研究AI辅助与绩效反馈如何影响专家决策。第一项研究中,医生先独立做出诊断(人类),再查看AI预测,最后调整为最终判断(人机协作);第二项研究(经记忆清除期后)则让同一批参与者先了解第一项研究中的聚合性能数据(包括自身、AI及人机团队的表现),再直接查看AI预测并立即做出诊断(无独立初始判断)。这两种流程模拟了临床中可能的实际使用方式,其中第二项模拟医生可依据过往表现调整对AI的信任度。结果显示,人机团队始终优于人类单独判断,但仍逊于纯AI,原因在于医生存在低估依赖现象,与以往众包工作者研究一致。提供绩效反馈并未显著改善人机团队表现,尽管提前展示AI结果会促使医生更倾向于采纳其建议。同时发现,人机团队的集成表现可超越纯AI,提示了未来人机协作的优化方向。

原文摘要 · Abstract (English)

Despite the growing interest in human-AI decision making, experimental studies with domain experts remain rare, largely due to the complexity of working with domain experts and the challenges in setting up realistic experiments. In this work, we conduct an in-depth collaboration with radiologists in prostate cancer diagnosis based on MRI images. Building on existing tools for teaching prostate cancer diagnosis, we develop an interface and conduct two experiments to study how AI assistance and performance feedback shape the decision making of domain experts. In Study 1, clinicians were asked to provide an initial diagnosis (human), then view the AI's prediction, and subsequently finalize their decision (human-AI team). In Study 2 (after a memory wash-out period), the same participants first received aggregated performance statistics from Study 1, specifically their own performance, the AI's performance, and their human-AI team performance, and then directly viewed the AI's prediction before making their diagnosis (i.e., no independent initial diagnosis). These two workflows represent realistic ways that clinical AI tools might be used in practice, where the second study simulates a scenario where doctors can adjust their reliance and trust on AI based on prior performance feedback. Our findings show that, while human-AI teams consistently outperform humans alone, they still underperform the AI due to under-reliance, similar to prior studies with crowdworkers. Providing clinicians with performance feedback did not significantly improve the performance of human-AI teams, although showing AI decisions in advance nudges people to follow AI more. Meanwhile, we observe that the ensemble of human-AI teams can outperform AI alone, suggesting promising directions for human-AI collaboration.

AI医疗人机协作影像诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。