用模拟医学生口试的方式,测试大模型在不确定中逐步推理的临床判断能力。
Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models
- 设计多轮交互式病例,让模型像医生一样主动追问和选检查
- 模型在不确定性下诊断准确率显著下降,暴露思维缺陷
- 适合评估医疗AI、研究智能体决策,尤其关注临床推理
医学临床推理是医生基于有限信息,通过针对性问诊、查体和检验逐步修正诊断的假设驱动过程。当前大语言模型(LLMs)的医学测评多为单轮问答,一次性给出完整病史信息,难以反映真实诊疗逻辑。为此,我们提出VivaBench,一个支持多轮交互的基准,用于评估LLM代理在序列化临床推理中的表现。数据集包含1762个由医师精心设计的临床情景,模拟医学生训练中的口试环节,要求模型主动探查关键发现、选择恰当检查,并在多步中整合信息作出诊断。尽管现有模型在完整描述病例时表现良好,但在需要迭代推理与应对不确定性时性能大幅下降。分析揭示了与临床实践相似的错误模式:(1) 过早锁定初始假设,(2) 检查顺序不当,(3) 未充分排查即过早定论,(4) 忽视危重疾病筛查。这些现象暴露出当前大模型在不确定环境下推理与决策的根本局限。VivaBench为评估对话式医疗AI系统提供了标准化框架,同时也为智能体在复杂决策场景中的行为研究贡献新视角。
原文摘要 · Abstract (English)
Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks for large language models (LLMs) primarily assess knowledge recall through single-turn questions, where complete clinical information is provided upfront. To address this gap, we introduce VivaBench, a multi-turn benchmark that evaluates sequential clinical reasoning in LLM agents. Our dataset consists of 1762 physician-curated clinical vignettes structured as interactive scenarios that simulate a (oral) examination in medical training, requiring agents to actively probe for relevant findings, select appropriate investigations, and synthesize information across multiple steps to reach a diagnosis. While current LLMs demonstrate competence in diagnosing conditions from well-described clinical presentations, their performance degrades significantly when required to navigate iterative diagnostic reasoning under uncertainty in our evaluation. Our analysis identified several failure modes that mirror common cognitive errors in clinical practice, including: (1) fixation on initial hypotheses, (2) inappropriate investigation ordering, (3) premature diagnostic closure, and (4) failing to screen for critical conditions. These patterns reveal fundamental limitations in how current LLMs reason and make decisions under uncertainty. Through VivaBench, we provide a standardized benchmark for evaluating conversational medical AI systems for real-world clinical decision support. Beyond medical applications, we contribute to the larger corpus of research on agentic AI by demonstrating how sequential reasoning trajectories can diverge in complex decision-making environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。