arXiv:2507.07988cs.CL2025-07被引 30

构建可解释的医疗推理评估基准,让大模型表现更可信。

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

  • 用专家级分步推理题构建500道医学难题库
  • 通过大模型自评机制实现与专家判断高度一致
  • 小模型反而超过部分大厂闭源模型,适合临床应用

随着大型语言模型(LLMs)逐步融入临床决策,确保其推理过程透明可信至关重要。然而现有评估方法或准确性不足,或难以扩展,缺乏严谨的基准。为此,我们提出MedThink-Bench,一个涵盖十个医学领域的500道高难度问题基准,每题均配有专家撰写的分步推理过程。基于此,我们设计了LLM-w-Ref评估框架,利用细粒度推理链与大模型作为裁判机制,在保持可扩展性的同时实现专家级推理评估。实验表明,该框架与专家判断具有强正相关性。对十二个前沿大模型的评测显示,较小模型(如MedGemma-27B)甚至优于部分大型闭源模型(如OpenAI-o3)。MedThink-Bench为评估大模型医疗推理能力提供了基础工具,推动其在临床实践中的安全落地。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasoning capability either suffer from unsatisfactory assessment or poor scalability, and a rigorous benchmark remains lacking. To address this, we introduce MedThink-Bench, a benchmark designed for rigorous, explainable, and scalable assessment of LLMs' medical reasoning. MedThink-Bench comprises 500 challenging questions across ten medical domains, each annotated with expert-crafted step-by-step rationales. Building on this, we propose LLM-w-Ref, a novel evaluation framework that leverages fine-grained rationales and LLM-as-a-Judge mechanisms to assess intermediate reasoning with expert-level fidelity while maintaining scalability. Experiments show that LLM-w-Ref exhibits a strong positive correlation with expert judgments. Benchmarking twelve state-of-the-art LLMs, we find that smaller models (e.g., MedGemma-27B) can surpass larger proprietary counterparts (e.g., OpenAI-o3). Overall, MedThink-Bench offers a foundational tool for evaluating LLMs' medical reasoning, advancing their safe and responsible deployment in clinical practice.

医疗AI大模型评估可解释性临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。