arXiv:2508.21430cs.CLcs.AI2025-08

首个专用于医疗多模态大模型奖励模型的评估基准

Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models

  • 构建覆盖13个器官系统、8个科室的多模态医疗数据集
  • 通过专家标注的1026例病例评估模型临床判断准确性
  • 发现现有模型在专业对齐上存在显著偏差,适合医学AI研究者

多模态大语言模型(MLLMs)在疾病诊断和临床决策等医疗应用中潜力巨大。然而,这些任务需要高度准确、上下文敏感且符合专业规范的输出,因此可靠的奖励模型和评判工具至关重要。尽管如此,医疗奖励模型(MRMs)和评判标准仍缺乏系统研究,尚无专门针对临床需求的评估基准。现有基准多聚焦通用MLLM能力或将其作为求解器评价,忽视了诊断准确性和临床相关性等关键维度。为此,我们提出Med-RewardBench,首个专为医疗场景设计的MRMs与评判模型评估基准。该基准包含覆盖13个器官系统和8个临床科室的多模态数据集,共1,026例专家标注案例,并采用三步严格流程确保数据质量,涵盖六个临床关键维度。我们评估了32个顶尖MLLMs(包括开源、商业及医疗专用模型),发现其输出与专家判断存在显著偏差。此外,我们还开发了基线模型,通过微调实现了显著性能提升。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) hold significant potential in medical applications, including disease diagnosis and clinical decision-making. However, these tasks require highly accurate, context-sensitive, and professionally aligned responses, making reliable reward models and judges critical. Despite their importance, medical reward models (MRMs) and judges remain underexplored, with no dedicated benchmarks addressing clinical requirements. Existing benchmarks focus on general MLLM capabilities or evaluate models as solvers, neglecting essential evaluation dimensions like diagnostic accuracy and clinical relevance. To address this, we introduce Med-RewardBench, the first benchmark specifically designed to evaluate MRMs and judges in medical scenarios. Med-RewardBench features a multimodal dataset spanning 13 organ systems and 8 clinical departments, with 1,026 expert-annotated cases. A rigorous three-step process ensures high-quality evaluation data across six clinically critical dimensions. We evaluate 32 state-of-the-art MLLMs, including open-source, proprietary, and medical-specific models, revealing substantial challenges in aligning outputs with expert judgment. Additionally, we develop baseline models that demonstrate substantial performance improvements through fine-tuning.

医疗AI多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。