评测大模型在真实临床案例中的推理能力,发现其诊断准确率超85%但关键步骤常遗漏。
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
- 构建1453例结构化病历数据集,覆盖13个系统10个专科,含罕见病。
- 提出自动评分框架,对推理过程的效率、真实性、完整性进行客观评估。
- 开源模型如DeepSeek-R1已逼近闭源系统,推动医疗AI普惠发展。
近期增强推理能力的大语言模型(如DeepSeek-R1和OpenAI-o3)取得显著进展,但在专业医疗场景中的应用仍缺乏深入探索,尤其缺乏对推理过程质量与最终输出的综合评估。为此,我们提出MedR-Bench,一个包含1,453例结构化患者病例的数据集,其标注依据临床病例报告生成推理参考。涵盖13个身体系统和10个专科,包含常见及罕见疾病。为全面评估模型表现,我们设计了三阶段评估框架:检查建议、诊断决策和治疗计划,模拟完整的患者诊疗流程。针对推理质量评估,我们开发了新型自动化系统Reasoning Evaluator,通过动态交叉验证与证据核查,对自由文本推理内容进行效率、实际性与完整性的客观评分。基于该基准,我们评估了五种前沿推理型LLM,包括DeepSeek-R1、OpenAI-o3-mini和Gemini-2.0-Flash Thinking等。结果显示,当提供充分检查结果时,当前模型在简单诊断任务中准确率超过85%;但在复杂任务如检查推荐和治疗规划中性能下降。尽管推理输出整体可靠(事实性得分超90%),但关键推理步骤常被遗漏。这些发现揭示了临床大模型的进步与局限。值得注意的是,开源模型如DeepSeek-R1正快速缩小与闭源系统的差距,凸显其在推动医疗AI可及性与公平性方面的潜力。
原文摘要 · Abstract (English)
Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in professional medical contexts remains underexplored, particularly in evaluating the quality of their reasoning processes alongside final outputs. Here, we introduce MedR-Bench, a benchmarking dataset of 1,453 structured patient cases, annotated with reasoning references derived from clinical case reports. Spanning 13 body systems and 10 specialties, it includes both common and rare diseases. To comprehensively evaluate LLM performance, we propose a framework encompassing three critical examination recommendation, diagnostic decision-making, and treatment planning, simulating the entire patient care journey. To assess reasoning quality, we present the Reasoning Evaluator, a novel automated system that objectively scores free-text reasoning responses based on efficiency, actuality, and completeness using dynamic cross-referencing and evidence checks. Using this benchmark, we evaluate five state-of-the-art reasoning LLMs, including DeepSeek-R1, OpenAI-o3-mini, and Gemini-2.0-Flash Thinking, etc. Our results show that current LLMs achieve over 85% accuracy in relatively simple diagnostic tasks when provided with sufficient examination results. However, performance declines in more complex tasks, such as examination recommendation and treatment planning. While reasoning outputs are generally reliable, with factuality scores exceeding 90%, critical reasoning steps are frequently missed. These findings underscore both the progress and limitations of clinical LLMs. Notably, open-source models like DeepSeek-R1 are narrowing the gap with proprietary systems, highlighting their potential to drive accessible and equitable advancements in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。