arXiv:2505.11413cs.CL2025-05NeurIPS被引 33

构建医疗大模型安全评估基准,测试其抗攻击与拒答能力

CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs

  • 设计涵盖8类医疗安全原则的1.8万条提示,分4级危害与4种攻击形式
  • 发现主流模型易被巧妙改写指令绕过防御,且误拒正常提问
  • 提出轻量检测器+提醒引导策略,有效提升模型安全响应

大型语言模型在医疗领域应用日益广泛,但其安全性、对齐性及对抗攻击脆弱性引发关注。现有基准缺乏临床针对性、危害等级划分和对越狱攻击的覆盖。本文提出CARES(临床对抗鲁棒性与安全评估)基准,包含超过18,000条提示,覆盖八项医疗安全原则、四级危害程度及四种提示风格:直接、间接、混淆和角色扮演,模拟恶意与非恶意使用场景。我们提出三类响应评估标准(接受、谨慎、拒绝)和细粒度安全评分,分析显示多个前沿模型仍易受微调指令绕过,且过度拒绝表述异常的正常查询。为此,我们提出基于轻量分类器的检测方法,通过提醒式条件化引导模型更安全响应。CARES为医疗LLM在对抗与模糊情境下的安全测试与改进提供严谨框架。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded harmfulness levels, and coverage of jailbreak-style attacks. We introduce CARES (Clinical Adversarial Robustness and Evaluation of Safety), a benchmark for evaluating LLM safety in healthcare. CARES includes over 18,000 prompts spanning eight medical safety principles, four harm levels, and four prompting styles: direct, indirect, obfuscated, and role-play, to simulate both malicious and benign use cases. We propose a three-way response evaluation protocol (Accept, Caution, Refuse) and a fine-grained Safety Score metric to assess model behavior. Our analysis reveals that many state-of-the-art LLMs remain vulnerable to jailbreaks that subtly rephrase harmful prompts, while also over-refusing safe but atypically phrased queries. Finally, we propose a mitigation strategy using a lightweight classifier to detect jailbreak attempts and steer models toward safer behavior via reminder-based conditioning. CARES provides a rigorous framework for testing and improving medical LLM safety under adversarial and ambiguous conditions.

医疗大模型安全评估对抗鲁棒性越狱检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。