模型安全保护不均,少数群体易受攻击
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment
- 构建多语言对抗测试集,检测模型对不同群体的防护差异
- 同一模型对不同群体防御率差距达42%,暴露安全不平等
- 通过针对性优化实现跨群体安全泛化,适合安全研究者参考
当前大语言模型的安全评估通过将伤害归入如“身份仇恨”等通用类别,制造出全面保护的虚假印象,掩盖了对特定群体的脆弱性。本文揭示‘选择性安全陷阱’:模型对某些群体强防护,却使代表性不足的群体极易遭受相同攻击。为此,我们提出MiJaBench——一个包含43,961个受控越狱提示的中英双语对抗基准,覆盖16个少数群体。在14个前沿LLM上评估后,生成615,454组提示-响应数据(MiJaBench-Align),发现安全对齐并非统一语义能力,而是一种基于人口的层级结构,同一模型对不同群体的防御率波动高达42%。这种差异在不同架构与语言间持续存在,并随模型规模放大,表明现有对齐方法学习的是群体特定防护,而非通用危害概念。通过对10亿参数基线模型进行定向直接偏好优化(DPO),我们实现了对全新群体和复杂攻击策略的强零样本安全泛化。所有数据集与脚本已公开,为实现公平、可迁移的安全对齐提供路径。
原文摘要 · Abstract (English)
Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under generic categories such as "Identity Hate", obscuring vulnerabilities toward specific populations. In this work, we expose the Selective Safety Trap: a systemic failure mode where models robustly defend specific populations while leaving underrepresented communities highly vulnerable to identical adversarial attacks. To systematically audit this phenomenon, we introduce MiJaBench, a bilingual (English-Portuguese) adversarial benchmark comprising 43,961 controlled jailbreaking prompts across 16 minority groups. By evaluating 14 state-of-the-art LLMs on MiJaBench, we curate 615,454 prompt-response pairs that compose MiJaBench-Align, revealing that safety alignment is not a uniform semantic capability but a demographic hierarchy, with defense rates fluctuating by up to 42% within the same model solely based on the target group. This disparity persists across architectures and languages and is amplified by scaling, indicating that current alignment methods learn group-specific safeguards rather than a generalized notion of harm. Through targeted direct preference optimization (DPO) on a 1B-parameter baseline, we achieve strong zero-shot safety generalizations to entirely unseen demographics and complex attack strategies. We release all datasets and scripts to provide the community with a concrete pathway toward equitable, transferable safety alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。