四款主流大模型在医疗问答中存在安全隐患,最高不当回答率达43.2%。
Large language models provide unsafe answers to patient-posed medical questions
- 用真实患者提问数据集对比四大模型安全表现
- 不当回答率21.6%至43.2%,严重错误达5%至13%
- 适合关注AI医疗安全的临床医生与政策制定者
数百万患者已定期使用大语言模型(LLM)聊天机器人获取医疗建议,引发患者安全担忧。本研究由医师主导,采用红队测试方法,对四个公开可用的聊天机器人——Anthropic的Claude、Google的Gemini、OpenAI的GPT-4o和Meta的Llama3-70B——在新构建的HealthAdvice数据集上进行评估。共分析了222个初级保健领域的患者提问,涵盖内科、妇产科和儿科内容,总计评估888条聊天机器人回复。结果显示各模型间存在显著差异:不当回答率从Claude的21.6%到Llama的43.2%不等,其中存在潜在严重伤害风险的回应占比为Claude的5%至GPT-4o和Llama的13%。定性分析发现部分回复可能造成严重患者伤害。研究提示,数百万患者可能正接收来自公开聊天机器人的不安全医疗建议,亟需提升这些强大工具的临床安全性。
原文摘要 · Abstract (English)
Millions of patients are already using large language model (LLM) chatbots for medical advice on a regular basis, raising patient safety concerns. This physician-led red-teaming study compares the safety of four publicly available chatbots--Claude by Anthropic, Gemini by Google, GPT-4o by OpenAI, and Llama3-70B by Meta--on a new dataset, HealthAdvice, using an evaluation framework that enables quantitative and qualitative analysis. In total, 888 chatbot responses are evaluated for 222 patient-posed advice-seeking medical questions on primary care topics spanning internal medicine, women's health, and pediatrics. We find statistically significant differences between chatbots. The rate of problematic responses varies from 21.6 percent (Claude) to 43.2 percent (Llama), with unsafe responses varying from 5 percent (Claude) to 13 percent (GPT-4o, Llama). Qualitative results reveal chatbot responses with the potential to lead to serious patient harm. This study suggests that millions of patients could be receiving unsafe medical advice from publicly available chatbots, and further work is needed to improve the clinical safety of these powerful tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。