评测大模型生成临床问诊模板的能力,发现其常冗长且忽略关键问题优先级。
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
- 用多代理流程优化提示并自动评分,评估模型生成结构化问诊问题的能力。
- o3模型最全面(达92.2%),但输出过长,且在长度限制下无法正确排序重点问题。
- 精神科、疼痛医学等叙事性强的领域表现显著下降,提示需更贴近临床实际的评估方法。
本研究评估大语言模型(LLMs)生成电子问诊结构化模板的能力。基于斯坦福eConsult团队开发并日常使用的145份专家模板,我们测试了前沿模型(包括o3、GPT-4o、Kimi K2、Claude 4 Sonnet、Llama 3 70B和Gemini 2.5 Pro)生成临床连贯、简洁且有优先级的问题框架的能力。通过结合提示优化、语义自动评分与优先级分析的多代理流程,结果表明:尽管o3模型在全面性上表现优异(最高达92.2%),但其输出普遍过长,且在长度受限条件下无法正确突出最关键临床问题。不同专科表现差异显著,在以叙事为主的领域如精神病学和疼痛医学中性能明显下降。研究证明LLM可提升医生间结构化信息交流效率,但强调亟需更稳健的评估方法,以衡量模型在真实临床沟通时间压力下对关键信息的优先处理能力。
原文摘要 · Abstract (English)
This study evaluates the capacity of large language models (LLMs) to generate structured clinical consultation templates for electronic consultation. Using 145 expert-crafted templates developed and routinely used by Stanford's eConsult team, we assess frontier models -- including o3, GPT-4o, Kimi K2, Claude 4 Sonnet, Llama 3 70B, and Gemini 2.5 Pro -- for their ability to produce clinically coherent, concise, and prioritized clinical question schemas. Through a multi-agent pipeline combining prompt optimization, semantic autograding, and prioritization analysis, we show that while models like o3 achieve high comprehensiveness (up to 92.2\%), they consistently generate excessively long templates and fail to correctly prioritize the most clinically important questions under length constraints. Performance varies across specialties, with significant degradation in narrative-driven fields such as psychiatry and pain medicine. Our findings demonstrate that LLMs can enhance structured clinical information exchange between physicians, while highlighting the need for more robust evaluation methods that capture a model's ability to prioritize clinically salient information within the time constraints of real-world physician communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。