评估12个大模型在临床笔记生成中的可靠性,发现小模型更稳定且符合隐私要求。
Are LLMs reliable? An exploration of the reliability of large language models in clinical note generation
- 对比12个开源与商用大模型在相同提示下的输出一致性
- 多数模型生成内容与专家笔记语义相近,Meta Llama 70B最可靠
- 推荐本地部署小型开源模型以保障数据隐私与效率
由于医疗提供者需对文档准确性及患者数据隐私承担法律责任,大语言模型(LLMs)响应的自然变异性给基于其的临床笔记生成(CNG)系统在真实临床流程中的应用带来挑战。CNG文本的复杂性进一步加剧了这一问题。为增强医疗提供者对LLM驱动工具的信心,本研究评估了来自Anthropic、Meta、Mistral和OpenAI的12个开源与专有LLMs在CNG中的可靠性,考察其在多次迭代中生成笔记的字符串一致性(一致率)、语义一致性(语义等价)和正确性(语义相似度)。结果表明:(1)所有模型家族均表现稳定,尽管表达方式不同,但语义保持一致;(2)多数模型生成的笔记与专家笔记语义接近。总体而言,Meta的Llama 70B表现最佳,其次为Mistral的小型模型。据此建议:将相对较小的开源模型本地部署于临床笔记生成,以满足数据隐私合规要求,并提升医疗提供者的文档效率。
原文摘要 · Abstract (English)
Due to the legal and ethical responsibilities of healthcare providers (HCPs) for accurate documentation and protection of patient data privacy, the natural variability in the responses of large language models (LLMs) presents challenges for incorporating clinical note generation (CNG) systems, driven by LLMs, into real-world clinical processes. The complexity is further amplified by the detailed nature of texts in CNG. To enhance the confidence of HCPs in tools powered by LLMs, this study evaluates the reliability of 12 open-weight and proprietary LLMs from Anthropic, Meta, Mistral, and OpenAI in CNG in terms of their ability to generate notes that are string equivalent (consistency rate), have the same meaning (semantic consistency) and are correct (semantic similarity), across several iterations using the same prompt. The results show that (1) LLMs from all model families are stable, such that their responses are semantically consistent despite being written in various ways, and (2) most of the LLMs generated notes close to the corresponding notes made by experts. Overall, Meta's Llama 70B was the most reliable, followed by Mistral's Small model. With these findings, we recommend the local deployment of these relatively smaller open-weight models for CNG to ensure compliance with data privacy regulations, as well as to improve the efficiency of HCPs in clinical documentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。