arXiv:2512.02024cs.CLcs.CY2025-12综述

构建临床处方审查基准,验证顶尖大模型已超人类药师水平

Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review

  • 设计涵盖14类常见处方错误的RxBench基准测试集
  • 三款顶级模型在准确率与鲁棒性上均超越其他模型和部分药师
  • 通过针对性微调使普通模型达到顶尖水平,适合医疗AI研发者使用

大语言模型(LLMs)在临床决策支持中的应用迅速发展,尤其体现在处方审查领域。为实现系统化、细粒度评估,我们构建了RxBench——一个覆盖常见处方审查类别并整合14种高频处方错误的综合性基准,错误来源为权威药学参考文献。RxBench包含1,150道单选题、230道多选题和879道简答题,所有题目均由经验丰富的临床药师审核。我们对18个前沿大模型进行了基准测试,发现其性能存在明显分层。值得注意的是,Gemini-2.5-pro-preview-05-06、Grok-4-0709和DeepSeek-R1-0528 consistently处于第一梯队,在准确率与鲁棒性上均优于其他模型。与持证药师对比显示,领先模型在某些任务中可达到甚至超过人类表现。此外,基于基准分析结果,我们对中等水平模型进行针对性微调,生成的专用模型在简答题任务上的表现已接近顶级通用大模型。RxBench的核心贡献在于建立了一个以错误类型为导向的标准化评估框架,不仅揭示了前沿大模型在处方审查中的能力与局限,也为开发更可靠、更专业的临床工具提供了基础资源。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has accelerated their integration into clinical decision support, particularly in prescription review. To enable systematic and fine-grained evaluation, we developed RxBench, a comprehensive benchmark that covers common prescription review categories and consolidates 14 frequent types of prescription errors drawn from authoritative pharmacy references. RxBench consists of 1,150 single-choice, 230 multiple-choice, and 879 short-answer items, all reviewed by experienced clinical pharmacists. We benchmarked 18 state-of-the-art LLMs and identified clear stratification of performance across tasks. Notably, Gemini-2.5-pro-preview-05-06, Grok-4-0709, and DeepSeek-R1-0528 consistently formed the first tier, outperforming other models in both accuracy and robustness. Comparisons with licensed pharmacists indicated that leading LLMs can match or exceed human performance in certain tasks. Furthermore, building on insights from our benchmark evaluation, we performed targeted fine-tuning on a mid-tier model, resulting in a specialized model that rivals leading general-purpose LLMs in performance on short-answer question tasks. The main contribution of RxBench lies in establishing a standardized, error-type-oriented framework that not only reveals the capabilities and limitations of frontier LLMs in prescription review but also provides a foundational resource for building more reliable and specialized clinical tools.

大模型评测医疗AI处方审查临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。