邀请临床医生与工程师协作,测试大模型在医疗场景下的安全漏洞。
Red Teaming Large Language Models for Healthcare
- 联合临床与计算专家开展红队测试,模拟真实医疗场景下的风险提示。
- 发现多类可能导致临床危害的模型漏洞,涵盖诊断、用药等关键环节。
- 结果可复现,适用于评估医疗大模型的安全性,适合研发与审核团队参考。
我们报告了2024年机器学习与医疗会议(ML4HC)会前研讨会‘针对医疗领域的大语言模型红队测试’的设计过程与发现,该活动于2024年8月15日举行。参会者包括兼具计算与临床背景的专家,尝试通过设计可能引发临床危害的真实医疗提示,来发现大语言模型(LLM)的潜在漏洞。红队测试结合临床医生参与,识别出开发人员因缺乏临床知识而难以察觉的模型风险。我们总结了所发现的漏洞并进行分类,并对所有提供的大语言模型进行了复制研究,验证其脆弱性表现。
原文摘要 · Abstract (English)
We present the design process and findings of the pre-conference workshop at the Machine Learning for Healthcare Conference (2024) entitled Red Teaming Large Language Models for Healthcare, which took place on August 15, 2024. Conference participants, comprising a mix of computational and clinical expertise, attempted to discover vulnerabilities -- realistic clinical prompts for which a large language model (LLM) outputs a response that could cause clinical harm. Red-teaming with clinicians enables the identification of LLM vulnerabilities that may not be recognised by LLM developers lacking clinical expertise. We report the vulnerabilities found, categorise them, and present the results of a replication study assessing the vulnerabilities across all LLMs provided.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。