arXiv:2509.16584cs.CLcs.AI2025-09EMNLP被引 12

诊断大模型医疗计算缺陷,提升临床可信度。

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

  • 分步评估公式选择、实体提取和计算,揭示隐藏错误
  • GPT-4o准确率从62.7%降至43.6%,暴露评估盲区
  • 无需微调的MedRaC系统可将准确率提至53.19%

大型语言模型(LLMs)在医学基准测试中表现优异,但其在医疗计算——临床决策的关键环节——的能力仍缺乏深入探索与严谨评估。现有基准常仅以宽泛数值容差判断最终答案,忽略系统性推理失误,可能引发严重临床误判。本文重新审视医疗计算评估,强调临床可信度。首先,我们清洗并重构MedCalc-Bench数据集,提出新的分步评估流程,独立评估公式选择、实体提取与算术计算。在此细粒度框架下,GPT-4o准确率从62.7%降至43.6%,暴露出先前评估掩盖的错误。其次,提出自动错误分析框架,生成结构化归因,经人工评估验证与专家判断高度一致,实现可扩展、可解释的诊断。最后,设计模块化代理管道MedRaC,结合检索增强生成与基于Python的代码执行。无需微调,即可将不同LLM的准确率从16.35%提升至53.19%。本工作揭示当前评估方法局限,提出更贴近临床的评测范式,推动构建可信赖的医学LLM应用。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical decision-making, remains underexplored and poorly evaluated. Existing benchmarks often assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. In this work, we revisit medical calculation evaluation with a stronger focus on clinical trustworthiness. First, we clean and restructure the MedCalc-Bench dataset and propose a new step-by-step evaluation pipeline that independently assesses formula selection, entity extraction, and arithmetic computation. Under this granular framework, the accuracy of GPT-4o drops from 62.7% to 43.6%, revealing errors masked by prior evaluations. Second, we introduce an automatic error analysis framework that generates structured attribution for each failure mode. Human evaluation confirms its alignment with expert judgment, enabling scalable and explainable diagnostics. Finally, we propose a modular agentic pipeline, MedRaC, that combines retrieval-augmented generation and Python-based code execution. Without any fine-tuning, MedRaC improves the accuracy of different LLMs from 16.35% up to 53.19%. Our work highlights the limitations of current benchmark practices and proposes a more clinically faithful methodology. By enabling transparent and transferable reasoning evaluation, we move closer to making LLM-based systems trustworthy for real-world medical applications.

医疗计算大模型评估推理诊断医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。