arXiv:2511.21701cs.CLcs.LG2025-11

小模型混合作专家架构在中医考试中超越大模型,打破规模迷信。

47B Mixture-of-Experts Beats 671B Dense Models on Chinese Medical Examinations

  • 采用混合专家架构,用更小模型实现高精度医疗问答。
  • 74.25%准确率超越671B大模型的64.07%,关键领域表现优异。
  • 适合医学教育与临床辅助系统研发人员参考。

大型语言模型(LLMs)在医疗领域的应用日益受到关注。本文对27个前沿LLM在中文医学考试题上的表现进行了全面基准评估,涵盖心血管、消化内科、血液科、传染病、肾病、神经科和呼吸科七个专业方向,覆盖主治医师与副主任医师两个级别。我们构建了一个包含2800道精心筛选题目的评估框架。结果显示,模型性能差异显著:Mixtral-8x7B以74.25%准确率位居榜首,优于DeepSeek-R1-671B的64.07%。值得注意的是,模型规模与性能无稳定相关性,小型混合作专家架构表现突出。不同专科间存在明显差距,心血管与神经科表现较好,而消化与肾病领域较弱。顶尖模型在主治与副高层级间性能下降极小,显示良好泛化能力。该基准为医学教育与临床决策支持系统部署提供关键洞见,揭示了当前技术在专业医疗场景中的潜力与局限。

原文摘要 · Abstract (English)

The rapid advancement of large language models(LLMs) has prompted significant interest in their potential applications in medical domains. This paper presents a comprehensive benchmark evaluation of 27 state-of-the-art LLMs on Chinese medical examination questions, encompassing seven medical specialties across two professional levels. We introduce a robust evaluation framework that assesses model performance on 2,800 carefully curated questions from cardiovascular, gastroenterology, hematology, infectious diseases, nephrology, neurology, and respiratory medicine domains. Our dataset distinguishes between attending physician and senior physician difficulty levels, providing nuanced insights into model capabilities across varying complexity. Our empirical analysis reveals substantial performance variations among models, with Mixtral-8x7B achieving the highest overall accuracy of 74.25%, followed by DeepSeek-R1-671B at 64.07%. Notably, we observe no consistent correlation between model size and performance, as evidenced by the strong performance of smaller mixture-of-experts architectures. The evaluation demonstrates significant performance gaps between medical specialties, with models generally performing better on cardiovascular and neurology questions compared to gastroenterology and nephrology domains. Furthermore, our analysis indicates minimal performance degradation between attending and senior physician levels for top-performing models, suggesting robust generalization capabilities. This benchmark provides critical insights for the deployment of LLMs in medical education and clinical decision support systems, highlighting both the promise and current limitations of these technologies in specialized medical contexts.

医学AI专家模型中文评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。