MedBench v4评估中文医疗大模型、多模态模型与智能体,揭示安全短板并验证智能体架构优势。
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
- 构建超70万任务的全国性云平台,覆盖24主类91副类医学专科,分设LLM、多模态、智能体赛道。
- 基础模型平均得分54.1/100(最佳Claude Sonnet 4.5达62.5),安全伦理仅18.4/100;智能体平均达79.8/100。
- 首次证明治理意识驱动的智能体可显著提升临床闭环能力,尤其在安全任务上达88.9/100。
医疗大模型、多模态模型与智能体的快速发展亟需能反映真实临床流程与安全约束的评估框架。我们提出MedBench v4,一个覆盖超过70万项专家标注任务的全国性云基准测试平台,涵盖24个主要医学专科和91个次要专科,设有针对大语言模型、多模态模型及智能体的独立评测赛道。所有题目经多轮临床专家审核(来自500余家医疗机构),开放式回答由经人类评分校准的LLM-as-a-judge打分。我们评估了15个前沿模型:基础大模型平均得分为54.1/100(最高为Claude Sonnet 4.5,达62.5/100),但安全与伦理表现偏低(仅18.4/100)。多模态模型整体表现更差(平均47.5/100;最高为GPT-5,54.9/100),感知能力强但跨模态推理弱。基于相同底座构建的智能体显著提升端到端性能(平均79.8/100),其中基于Claude Sonnet 4.5的智能体总体得分高达85.3/100,安全任务达88.9/100。该结果揭示了基础模型在多模态推理与安全性方面的持续不足,同时表明具备治理意识的智能体编排可大幅增强临床就绪度,且不牺牲能力。通过对齐中国临床指南与监管优先事项,本平台为医院、开发者与政策制定者提供实用的医疗AI审计参考。
原文摘要 · Abstract (English)
Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud-based benchmarking infrastructure comprising over 700,000 expert-curated tasks spanning 24 primary and 91 secondary specialties, with dedicated tracks for LLMs, multimodal models, and agents. Items undergo multi-stage refinement and multi-round review by clinicians from more than 500 institutions, and open-ended responses are scored by an LLM-as-a-judge calibrated to human ratings. We evaluate 15 frontier models. Base LLMs reach a mean overall score of 54.1/100 (best: Claude Sonnet 4.5, 62.5/100), but safety and ethics remain low (18.4/100). Multimodal models perform worse overall (mean 47.5/100; best: GPT-5, 54.9/100), with solid perception yet weaker cross-modal reasoning. Agents built on the same backbones substantially improve end-to-end performance (mean 79.8/100), with Claude Sonnet 4.5-based agents achieving up to 85.3/100 overall and 88.9/100 on safety tasks. MedBench v4 thus reveals persisting gaps in multimodal reasoning and safety for base models, while showing that governance-aware agentic orchestration can markedly enhance benchmarked clinical readiness without sacrificing capability. By aligning tasks with Chinese clinical guidelines and regulatory priorities, the platform offers a practical reference for hospitals, developers, and policymakers auditing medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。