arXiv:2607.02879cs.AI2026-07

构建复杂医疗计算新基准,提升大模型在真实临床场景的推理能力。

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

论文配图:MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents
图 1 · 摘自论文原文
  • 设计多工具协同与嵌套调用的智能体框架,支持复杂医疗计算。
  • 在2268个真实病例上验证,跨三种任务设置均优于现有方法。
  • 适合医疗AI研究者和临床决策系统开发者参考使用。

当前医学计算评测基准多基于简化场景,每例患者对应单一计算器且工具明确指定。然而真实临床中常需多个计算器联合评估、嵌套式计算及模糊查询。为此,我们提出新基准MedCalc-Pro,涵盖单计算器、多计算器、嵌套计算器三类渐进挑战任务,包含2268个真实临床案例,覆盖14个科室的77种医学计算器。同时,提出通用性更强的智能体框架,支持多工具选择与嵌套调用,并通过结构化验证与证据审查抑制参数误差传播。系统比较开源、闭源及医疗专用大模型,结果表明该框架在三类任务中均表现最佳。本工作为复杂医学计算场景下的大模型评估与应用提供了新基准与方法。

原文摘要 · Abstract (English)

Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each patient case corresponds to a single calculator and the required tool is explicitly specified in the query. However, real clinical scenarios often require multiple calculators for joint evaluation, nested-scale calculation, and fuzzy queries that do not directly specify the target calculator. To this end, we propose a new medical calculation benchmark, MedCalc-Pro, which covers three progressively challenging task settings: single-calculator, multi-calculator, and nested-calculator calculation settings. MedCalc-Pro contains 2,268 real-world clinical cases, covering 77 medical calculators across 14 clinical departments. Meanwhile, to address the limited performance of existing frameworks and methods in complex clinical scenarios, we further propose a more generalizable agent framework that supports multi-tool selection and nested-tool calling, while suppressing parameter error propagation through structured validation and evidence review. We conduct systematic comparisons across open-source, closed-source, and medical-specialized LLMs, and the results show that our framework achieves the best performance across all three task settings. This work provides a new benchmark and method for evaluating and applying LLMs in challenging medical calculation scenarios.

医疗AI大模型智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。