arXiv:2506.02442cs.CL2025-06被引 8

模型解密能力可能让安全机制失效,引发危险输出或过度拒绝。

Should LLM Safety Be More Than Refusing Harmful Instructions?

  • 构建双维安全评估框架:拒绝对有害指令、生成内容是否安全。
  • 发现具备解密能力的模型在长尾加密文本中至少一维安全失效。
  • 适合关注大模型安全边界与对抗攻击防御的研究者阅读。

本文系统评估了大语言模型(LLMs)在长尾分布(加密)文本上的行为及其安全影响。提出二维安全评估框架:(1) 指令拒绝能力——拒绝对有害混淆指令;(2) 生成安全——抑制生成有害响应。通过全面实验发现,具备解密能力的模型可能面临不匹配泛化攻击:其安全机制至少在一个维度失效,导致不安全输出或过度拒绝。基于此,评估了多种预训练和后训练安全防护措施,讨论其优缺点。本研究深化了对长尾文本场景下大模型安全性的理解,并为构建鲁棒安全机制提供方向。

原文摘要 · Abstract (English)

This paper presents a systematic evaluation of Large Language Models' (LLMs) behavior on long-tail distributed (encrypted) texts and their safety implications. We introduce a two-dimensional framework for assessing LLM safety: (1) instruction refusal-the ability to reject harmful obfuscated instructions, and (2) generation safety-the suppression of generating harmful responses. Through comprehensive experiments, we demonstrate that models that possess capabilities to decrypt ciphers may be susceptible to mismatched-generalization attacks: their safety mechanisms fail on at least one safety dimension, leading to unsafe responses or over-refusal. Based on these findings, we evaluate a number of pre-LLM and post-LLM safeguards and discuss their strengths and limitations. This work contributes to understanding the safety of LLM in long-tail text scenarios and provides directions for developing robust safety mechanisms.

大模型安全加密文本对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。