arXiv:2601.11567cs.CLcs.AI2026-01中稿 · 47 workshop Reprod…

小模型在儿科内分泌领域表现不稳定,准确率高不等于可靠。

Measuring Stability Beyond Accuracy in Small Open-Source Medical Large Language Models for Pediatric Endocrinology

  • 用多选题+专家评审评估6个开源医学大模型,关注一致性与鲁棒性。
  • 高一致输出未必正确,部分模型存在自评偏差和解释顺序依赖。
  • 微小提示变化或系统环境差异就导致结果漂移,影响临床应用可靠性。

小型开源医学大语言模型(LLM)在低资源部署和普及方面具有潜力,但现有评估常局限于多选题(MCQ)准确率,缺乏对一致性、鲁棒性及推理行为的考察。本文结合多选题与人工评估、临床专家评审,对六种小型开源医学LLM(HuatuoGPT-o1、Diabetica-7B、Diabetica-o1、Meditron3-8B、MedFound-7B、ClinicaGPT-base-zh)在儿科内分泌领域的表现进行评估。在确定性设置下,分析提示变化对输出及自我评估偏差的影响;在随机性设置下,评估输出变异性,并探究一致性与正确性的关系。结果显示,HuatuoGPT-o1-8B表现最优,但高一致性并非正确性的指标。当要求选择正确推理路径时,HuatuoGPT-o1-8B与Diabetica-o1均表现出自评偏差和对候选解释顺序的依赖。专家评审发现错误推理理由中既有临床可接受内容,也存在临床疏漏。进一步表明,系统级扰动(如CUDA构建差异)虽不影响准确率,却可能导致统计显著的输出变化。研究揭示:语义细微的提示扰动即可引发输出分歧,挑战了基于LLM评估的可重复性,凸显不同随机环境下输出变异性问题,强调需建立更全面的诊断框架以理解真实临床决策支持中的潜在风险。

原文摘要 · Abstract (English)

Small open-source medical large language models (LLMs) offer promising opportunities for low-resource deployment and broader accessibility. However, their evaluation is often limited to accuracy on medical multiple choice question (MCQ) benchmarks, and lacks evaluation of consistency, robustness, or reasoning behavior. We use MCQ coupled to human evaluation and clinical review to assess six small open-source medical LLMs (HuatuoGPT-o1 (Chen 2024), Diabetica-7B, Diabetica-o1 (Wei 2024), Meditron3-8B (Sallinen2025), MedFound-7B (Liu 2025), and ClinicaGPT-base-zh (Wang 2023)) in pediatric endocrinology. In deterministic settings, we examine the effect of prompt variation on models' output and self-assessment bias. In stochastic settings, we evaluate output variability and investigate the relationship between consistency and correctness. HuatuoGPT-o1-8B achieved the highest performance. The results show that high consistency across the model response is not an indicator of correctness, although HuatuoGPT-o1-8B showed the highest consistency rate. When tasked with selecting correct reasoning, both HuatuoGPT-o1-8B and Diabetica-o1 exhibit self-assessment bias and dependency on the order of the candidate explanations. Expert review of incorrect reasoning rationales identified a mix of clinically acceptable responses and clinical oversight. We further show that system-level perturbations, such as differences in CUDA builds, can yield statistically significant shifts in model output despite stable accuracy. This work demonstrates that small, semantically negligible prompt perturbations lead to divergent outputs, raising concerns about reproducibility of LLM-based evaluations and highlights the output variability under different stochastic regimes, emphasizing the need of a broader diagnostic framework to understand potential pitfalls in real-world clinical decision support scenarios.

医学大模型稳定性评估儿科内分泌推理一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。