arXiv:2603.02222cs.LGcs.AI2026-03被引 1

MedCalc-Bench实测误判模型能力,开卷提示让准确率飙升至85%以上。

MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation

  • 开卷提示(提供计算公式)使模型准确率从52%升至81%-85%
  • 修复20多个关键错误后,原有高分结果多被推翻
  • 应改为工具使用评估,而非临床推理测试

MedCalc-Bench是评估大模型在临床计算任务中表现的常用基准,当前最优直接提示方法在验证集上准确率约35%(HELM MedHELM榜单),最佳已发表方法(带可验证奖励的强化学习)达74%。本文提出三项质疑:首先,系统审计发现该基准存在超过20处实现错误,包括关键公式偏差和运行时漏洞,部分源于NeurIPS论文数据集;其次,仅在推理时提供计算器说明(开卷提示),无需微调即可将GLM-4.6V和GLM-4.7的准确率从~52%提升至81%-85%,超越所有已发表成果;最后,利用GPT-5.2-Thinking建立95%-97%上限,残余误差主要来自标注错误与数据歧义。研究揭示该基准实际衡量的是公式记忆与算术精度,更适合作为工具使用评估。

原文摘要 · Abstract (English)

MedCalc-Bench is a widely used benchmark for evaluating LLM performance on clinical calculator tasks, with state-of-the-art direct prompting scores plateauing around 35% on the Verified split (HELM MedHELM leaderboard) and the best published approach-RL with verifiable rewards-reaching 74%. We present three contributions that challenge the benchmark's current framing. First, we conduct a systematic audit of the benchmark's calculator implementations, identifying and fixing over 20 errors ranging from critical formula inaccuracies to runtime bugs in a NeurIPS-published dataset. Second, we show that a simple intervention-providing the model with the calculator specification at inference time ("open-book" prompting)-raises accuracy from ~52% to 81-85% on GLM-4.6V and GLM-4.7, surpassing all published results including RL-trained systems, without any fine-tuning. Third, we establish an upper bound of 95-97% using GPT-5.2-Thinking, with residual errors attributable primarily to ground-truth issues and dataset ambiguities. Our findings suggest that MedCalc-Bench predominantly measures formula memorization and arithmetic precision rather than clinical reasoning, and would be better framed as a tool-use evaluation.

大模型评估医疗计算开卷提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。