arXiv:2604.04325cs.CL2026-04被引 4

测试大模型在多轮问诊中的诊断行为,发现提前下结论是主要问题。

Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction

  • 构建多轮医学诊断基准MINT,模拟真实临床推理过程。
  • 超55%的诊断在前两轮就定论,错误修正次数是正确变错的10.6倍。
  • 延迟关键信息暴露可提升准确率62.6%,避免诊断骤降23.3%。

大型语言模型(LLMs)在单轮提供全部临床信息时诊断准确率很高,但在更接近真实临床推理的多轮证据累积场景下表现尚不明确。我们提出MINT(Medical Incremental N-Turn Benchmark)——一个高保真、多轮医学诊断基准,包含1,035个病例,具有临床标注的证据片段、可控的回合粒度和信息保留的分解方式。对11个LLMs在MINT上的系统评估揭示了三种持续存在的行为模式:(1) 急于作答,模型在充分证据前即提交答案,超过55%的答案在前两轮内确定;(2) 自我修正,错误到正确的修改频率高达正确到错误的10.6倍,表明存在潜在自我纠错能力,但过早承诺会扼杀该能力;(3) 强诱因,如实验室结果等临床显著信息会引发提前回答,即使模型被明确要求等待。我们将这些发现转化为临床可操作建议:将诊断问题推迟至后续回合,可将首次承诺时的准确率提升62.6%;将显著临床证据延后呈现,可防止因过早承诺导致高达23.3%的准确率暴跌。本研究既提供了受控评估框架,也给出提升多轮医学诊断可靠性的重要建议。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve high accuracy in medical diagnosis when all clinical information is provided in a single turn, yet how they behave under multi-turn evidence accumulation closer to real clinical reasoning remains unexplored. We introduce MINT (Medical Incremental N-Turn Benchmark), a high-fidelity, multi-turn medical diagnosis benchmark comprising 1,035 cases with clinically labeled evidence shards, controlled turn granularity, and information-preserving decomposition. Through systematic evaluation of 11 LLMs on MINT, we uncover three persistent behavioral patterns that significantly impact diagnostic decisions: (1) intent to answer, models rush to answer before sufficient evidence has been observed, with over 55% of answers committed within the first two turns; (2) self-correction, incorrect-to-correct answer revisions occur at up to 10.6 times the rate of correct-to-incorrect flips, revealing a latent capacity for self-correction that premature commitment forecloses; and (3) strong lures, clinically salient information such as laboratory results trigger premature answering even when models are explicitly instructed to wait. We translate these findings into clinically actionable guidance: deferring the diagnostic question to later turns reduces premature answering and improves accuracy at the first point of commitment by up to 62.6%, while reserving salient clinical evidence for later turns prevents a catastrophic accuracy drop of up to 23.3% caused by premature commitment. Our work provides both a controlled evaluation framework and concrete recommendations for improving the reliability of LLMs in multi-turn medical diagnosis.

医疗AI大模型评估多轮对话诊断推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。