arXiv:2509.19344cs.CL2025-09被引 1

Llama3.1在重症医学问答中准确率达60%,70B版本比8B高30%。

Performance of Large Language Models in Answering Critical Care Medicine Questions

  • 用70亿参数模型测试重症医学问题,表现优于80亿参数版本。
  • 平均准确率60%,科研类问题最高(68.4%),肾脏类最低(47.9%)。
  • 适合关注临床智能辅助的医生与研究者,尤其需提升亚专科能力。

大型语言模型已在医学生水平的问题上进行测试,但在重症医学(CCM)等专业领域表现尚不明确。本研究评估了Meta-Llama 3.1模型(80亿和700亿参数)在871个重症医学问题上的表现。结果显示,Llama3.1:70B的准确率比8B高30%,平均准确率为60%。不同领域表现差异显著,科研类问题最高(68.4%),肾脏类最低(47.9%),表明未来需进一步拓展模型在各亚专科领域的覆盖能力。

原文摘要 · Abstract (English)

Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in Renal (47.9%), highlighting the need for broader future work to improve models across various subspecialty domains.

大模型重症医学医疗问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。