开源医学大模型评估套件,覆盖30项任务,可直接用于强化学习微调。
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

- 构建30个医学任务的开放评估套件,涵盖问答、信息抽取等
- 前沿模型(如Gemini 3 Pro)性能领先,且更省 token
- 支持后训练强化学习,适合医疗AI研发与评测人员使用
评估大语言模型在医疗应用中的表现仍面临基准饱和、数据获取受限及任务覆盖不足的问题。现有评估套件或已饱和、依赖封闭数据集,或缺乏全面模型覆盖。我们推出 Medmarks,一个完全开源的评估套件,包含30个基准测试,覆盖问答、信息抽取、医学计算和开放式临床推理。对61个模型、71种配置进行了系统性评估,采用可验证指标与LLM-as-a-Judge方法。结果显示:前沿推理模型(Gemini 3 Pro Preview、GPT-5.1、GPT-5.2)在各项基准上表现最佳;多数前沿专有模型比开源模型更高效;医学微调模型优于通用模型;小模型及Grok 4对答案顺序敏感。其中部分评估(Medmarks-T)可直接作为强化学习环境,用于后训练医学推理模型。代码已开源于 https://github.com/MedARC-AI/Medmarks。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either saturated, heavily depend on restricted datasets, or lack comprehensive model coverage. We introduce Medmarks, a fully open-source evaluation suite with 30 benchmarks spanning question answering, information extraction, medical calculations, and open-ended clinical reasoning. We perform a systematic evaluation of 61 models across 71 configurations using verifiable metrics and LLM-as-a-Judge. Our results show that frontier reasoning models (Gemini 3 Pro Preview, GPT-5.1, & GPT-5.2) achieve the highest performance across both benchmarks, most frontier proprietary models are significantly more token efficient than open-weight alternatives, medically fine-tuned models outperform their generalist counterparts, and that models are susceptible to answer-order bias (particularly smaller models and Grok 4). A subset of our evals (Medmarks-T) can be directly used as reinforcement learning environments to post-train LLMs for medical reasoning. Code is available at https://github.com/MedARC-AI/Medmarks
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。