arXiv:2511.20526cs.AI2025-11

测试大模型在中药师资格考中的表现,发现国产模型更优。

Assessing LLMs' Performance: Insights from the Chinese Pharmacist Exam

  • 用真实药考题测试两个大模型,仅用文字题
  • 深求-R1准确率90.0%,远超ChatGPT-4o的76.1%
  • 适合关注AI医疗评估和中文专业模型的研究者

随着大语言模型日益融入数字健康教育与测评流程,其在高风险、领域特定认证任务中的能力仍待深入探索。在中国,全国药师执业资格考试是评估药师临床与理论能力的标准化基准。本研究旨在对比ChatGPT-4o与DeepSeek-R1在2017–2021年真实药考题上的表现,并探讨其对AI辅助形成性评价的意义。共收集2,306道纯文本多选题,排除含表格或图像的题目。每题以原始中文输入,模型回答按完全正确率评估。采用皮尔逊卡方检验比较总体表现,费舍尔精确检验分析年度多选准确率。结果显示,DeepSeek-R1整体准确率达90.0%,显著高于ChatGPT-4o的76.1%(p < 0.001)。单元级分析显示,深求-R1在基础与临床综合模块均具持续优势。虽年度表现亦倾向深求-R1,但各年单元间差异均未达统计显著性(所有p > 0.05)。结论表明,深求-R1在结构与语义层面高度契合药考要求。该结果提示应进一步研究领域专用模型在此场景的应用潜力,同时强调在法律与伦理敏感情境下仍需人工监督。

原文摘要 · Abstract (English)

Background: As large language models (LLMs) become increasingly integrated into digital health education and assessment workflows, their capabilities in supporting high-stakes, domain-specific certification tasks remain underexplored.In China, the national pharmacist licensure exam serves as a standardized benchmark for evaluating pharmacists' clinical and theoretical competencies. Objective: This study aimed to compare the performance of two LLMs: ChatGPT-4o and DeepSeek-R1 on real questions from the Chinese Pharmacist Licensing Examination (2017-2021), and to discuss the implications of these performance differences for AI-enabled formative evaluation. Methods: A total of 2,306 multiple-choice (text-only) questions were compiled from official exams, training materials, and public databases. Questions containing tables or images were excluded. Each item was input in its original Chinese format, and model responses were evaluated for exact accuracy. Pearson's Chi-squared test was used to compare overall performance, and Fisher's exact test was applied to year-wise multiple-choice accuracy. Results: DeepSeek-R1 outperformed ChatGPT-4o with a significantly higher overall accuracy (90.0% vs. 76.1%, p < 0.001). Unit-level analyses revealed consistent advantages for DeepSeek-R1, particularly in foundational and clinical synthesis modules. While year-by-year multiple-choice performance also favored DeepSeek-R1, this performance gap did not reach statistical significance in any specific unit-year (all p > 0.05). Conclusion: DeepSeek-R1 demonstrated robust alignment with the structural and semantic demands of the pharmacist licensure exam. These findings suggest that domain-specific models warrant further investigation for this context, while also reinforcing the necessity of human oversight in legally and ethically sensitive contexts.

大模型评测医疗AI中文模型药考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。