arXiv:2605.05427cs.AI2026-05

21个大模型安全审计发现:拒绝对话不等于安全,保护不均且行为稳定。

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

论文配图:The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models
图 1 · 摘自论文原文
  • 通过调整数据组合,分离模型敏感度与数据毒性影响
  • 保守模型过度拒绝,宽松模型容忍更多有害响应;残障相关攻击防护最弱
  • 同一模型家族安全行为跨代稳定,说明训练目标比结构更重要

拒绝对话率并非大语言模型安全性的可靠指标:模型可能过度拒绝无害提示,同时仍会响应有害请求。我们在四个安全基准(OR-Bench、XSTest、ToxiGen、BOLD)上对21个开源大模型进行了大规模安全行为审计,采用组合调整方法以分离模型敏感度与数据毒性干扰。结果揭示三点:第一,不同模型生态采用根本不同的校准策略——如Llama等保守生态虽抑制不当输出但过拒率高,而DeepSeek、Qwen等宽松生态保持有用性但有害合规更高;第二,群体保护不均:模型过度保护主流种族和宗教群体,常拒绝关于他们的无害提问,但对残障相关攻击的防护明显不足;第三,同一模型家族在不同代际和规模下,拒绝对话与合规倾向保持稳定,表明后训练目标比模型架构更深刻地塑造安全行为。研究呼吁采用联合、具种族意识、多评估者视角的安全评测体系。

原文摘要 · Abstract (English)

Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both failure modes across 21 open-weight LLMs on four safety benchmarks (OR-Bench, XSTest, ToxiGen, BOLD), using a composition adjustment to isolate model sensitivity from dataset toxicity confounds. We report three findings. First, models adopt fundamentally different calibration strategies: conservative ecosystems such as Llama suppress unsafe outputs at the cost of elevated over-refusals, while permissive ecosystems such as DeepSeek and Qwen preserve helpfulness but tolerate higher harmful compliance. Second, demographic protection is unequal: models over-protect prominent racial and religious groups, frequently refusing even benign prompts about them, while providing substantially weaker protection against disability-targeted attacks. Third, refusal and compliance tendencies are stable within model families across generations and scales, suggesting that post-training objectives shape safety behavior more than architecture. Our results call for joint, demographically-aware, and multi-judge safety evaluation.

大模型安全拒绝机制公平性审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。