测试七款大模型在医疗场景下的越狱攻击脆弱性,发现安全风险极高。
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
- 用自动化评估框架测试三种越狱技术在医疗场景的攻击效果。
- 主流开源与商用模型均易被越狱,暴露严重安全隐患。
- 提出持续微调可提升防御能力,适合医疗AI安全研究者参考。
大型语言模型(LLMs)在医疗应用中日益普及,但其在临床实践中的部署引发重大安全担忧,包括有害信息传播风险。本研究系统评估了七种LLMs在医疗语境下对三种先进黑盒越狱技术的脆弱性。为量化攻击效果,我们提出一种自动化且领域适配的代理评估流程。实验结果表明,主流商业与开源模型在医疗越狱攻击面前高度脆弱。为进一步提升模型安全性与可靠性,我们还研究了持续微调(CFT)在抵御医疗对抗攻击中的有效性。研究强调需持续更新攻击评估方法、开展领域特定的安全对齐,并平衡模型安全与实用性。该工作为推进AI临床医生的安全性与可靠性提供了切实可行的洞见,助力医疗AI的伦理化与高效部署。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly utilized in healthcare applications. However, their deployment in clinical practice raises significant safety concerns, including the potential spread of harmful information. This study systematically assesses the vulnerabilities of seven LLMs to three advanced black-box jailbreaking techniques within medical contexts. To quantify the effectiveness of these techniques, we propose an automated and domain-adapted agentic evaluation pipeline. Experiment results indicate that leading commercial and open-source LLMs are highly vulnerable to medical jailbreaking attacks. To bolster model safety and reliability, we further investigate the effectiveness of Continual Fine-Tuning (CFT) in defending against medical adversarial attacks. Our findings underscore the necessity for evolving attack methods evaluation, domain-specific safety alignment, and LLM safety-utility balancing. This research offers actionable insights for advancing the safety and reliability of AI clinicians, contributing to ethical and effective AI deployment in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。