评估AI医生问诊能力,发现顶尖模型仍有明显短板
The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability
- 构建3000个模拟患者,覆盖语言、情绪、认知等真实行为差异
- 多维度评测显示顶尖模型在问诊效率与准确性上仍有巨大提升空间
- 揭示问诊策略需平衡诊断效果与实际临床场景的可行性
一位优秀的医生应具备共情力、专业能力、耐心和清晰沟通。近年来,AI医生已具备专家级诊断能力,尤其在主动提问获取信息方面取得进展,但其他关键素质仍被忽视。为弥补这一差距,我们提出MAQuE(Medical Agent Questioning Evaluation),这是首个大规模、自动化的医疗多轮问诊评估基准,包含3000个模拟患者代理,其具有多样化的语言模式、认知局限、情绪反应及被动披露倾向。我们还设计了多维度评估框架,涵盖任务成功率、问诊熟练度、对话能力、问诊效率和患者体验。在不同大模型上的实验表明,各模型在多个评估维度均面临显著挑战,即使顶尖模型在面对真实患者行为变化时也表现出高度敏感性,严重影响诊断准确率。细粒度指标进一步揭示了不同评估视角间的权衡关系,凸显了在真实临床环境中实现性能与实用性平衡的难度。
原文摘要 · Abstract (English)
An effective physician should possess a combination of empathy, expertise, patience, and clear communication when treating a patient. Recent advances have successfully endowed AI doctors with expert diagnostic skills, particularly the ability to actively seek information through inquiry. However, other essential qualities of a good doctor remain overlooked. To bridge this gap, we present MAQuE(Medical Agent Questioning Evaluation), the largest-ever benchmark for the automatic and comprehensive evaluation of medical multi-turn questioning. It features 3,000 realistically simulated patient agents that exhibit diverse linguistic patterns, cognitive limitations, emotional responses, and tendencies for passive disclosure. We also introduce a multi-faceted evaluation framework, covering task success, inquiry proficiency, dialogue competence, inquiry efficiency, and patient experience. Experiments on different LLMs reveal substantial challenges across the evaluation aspects. Even state-of-the-art models show significant room for improvement in their inquiry capabilities. These models are highly sensitive to variations in realistic patient behavior, which considerably impacts diagnostic accuracy. Furthermore, our fine-grained metrics expose trade-offs between different evaluation perspectives, highlighting the challenge of balancing performance and practicality in real-world clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。