医生输入影响AI诊断,专家意见提升准确率,错误引导则导致严重偏差。
Clinician input steers AI toward accurate and harmful recommendations
- 通过临床案例模拟医生推理,测试不同输入对AI诊断和建议的影响。
- 专家输入使诊断一致性从65.8%升至93.5%,错误引导使准确率平均下降5.4个百分点。
- 提出可缓解有害输出的交互式评估方法,适合医疗AI安全研究者使用。
大型语言模型(LLMs)正进入临床流程,但现有评估很少考察医生推理如何影响模型行为。基于61个精选的NEJM病例记录,我们测试了专家或误导性医生推理对21种不同推理变体(涵盖8个专有及开源模型)生成鉴别诊断与下一步建议的影响。在医生输入后,LLM与医生的一致性显著提升:具有≥3个重叠鉴别诊断的案例比例从65.8%升至93.5%,具有≥3个重叠建议的比例从20.3%升至53.8%。专家背景显著提升所有21个模型中正确最终诊断的包含率(均值+20.4个百分点),反映推理改进与被动内容复现;而对抗性背景在14个模型中显著降低性能(均值-5.4个百分点)。专家背景还显著提高所有模型的主要诊断准确性,对抗性背景则在13个模型中显著降低。多轮分歧挑战揭示了模型的不同特征,从高度顺从到固执己见,即使在较稳健的模型中,对抗性论点仍是弱点。推理时缩放将各世界卫生组织(WHO)危害等级中的有害复现减少62.7%(轻度)、57.9%(中度)、76.3%(重度)和83.5%(死亡级)。推理时提示策略在GPT-5、Claude Sonnet 4.5和Gemini 3 Flash上恢复了因对抗性输入损失的诊断准确率,同时保留专家输入优势,并大幅降低各危害等级下的高度一致有害复现。
原文摘要 · Abstract (English)
Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior during clinical interactions. Using 61 curated NEJM Case Records, we tested how expert or misleading clinician reasoning influenced AI-generated differential diagnoses and next step recommendations across 21 reasoning variants from 8 proprietary and open-source models. After clinician exposure, LLM-clinician concordance increased: simulations with >=3 overlapping differential diagnoses rose from 65.8% to 93.5%, and those with >=3 overlapping next step recommendations from 20.3% to 53.8%. Expert context significantly improved correct final-diagnosis inclusion in all 21 models (mean +20.4 pp), reflecting both improved reasoning and passive content echoing, while adversarial context significantly degraded performance in 14 models (mean -5.4 pp). Expert context also significantly increased leading-diagnosis accuracy in all 21 models, whereas adversarial context significantly reduced it in 13. Multi-turn disagreement challenges revealed distinct model phenotypes, from highly conformist to dogmatic, with adversarial arguments remaining a vulnerability even in otherwise resilient models. Inference-time scaling reduced harmful echoing of clinician-introduced recommendations across WHO harm-severity tiers by 62.7% for mild, 57.9% for moderate, 76.3% for severe, and 83.5% for death-tier recommendations. Inference-time prompting recovered diagnostic accuracy lost to adversarial context while preserving expert-context benefits across GPT-5, Claude Sonnet 4.5, and Gemini 3 Flash, and sharply reduced highly consistent harmful echoing across severity tiers. These findings provide a foundation for evaluating clinician-AI collaboration and introduce interactive metrics and mitigation strategies essential to safety and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。