Llama3.1在重症医学问答中准确率达60%,70B版本比8B高30%。
Performance of Large Language Models in Answering Critical Care Medicine Questions
- 用70亿参数模型测试重症医学问题,表现优于80亿参数版本。
- 平均准确率60%,科研类问题最高(68.4%),肾脏类最低(47.9%)。
- 适合关注临床智能辅助的医生与研究者,尤其需提升亚专科能力。
大型语言模型已在医学生水平的问题上进行测试,但在重症医学(CCM)等专业领域表现尚不明确。本研究评估了Meta-Llama 3.1模型(80亿和700亿参数)在871个重症医学问题上的表现。结果显示,Llama3.1:70B的准确率比8B高30%,平均准确率为60%。不同领域表现差异显著,科研类问题最高(68.4%),肾脏类最低(47.9%),表明未来需进一步拓展模型在各亚专科领域的覆盖能力。
原文摘要 · Abstract (English)
Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in Renal (47.9%), highlighting the need for broader future work to improve models across various subspecialty domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。