将有害请求转为多选题,可绕过大模型的安全拒绝机制。
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints

- 把有害内容包装成强制选择的多选题,让模型无法拒绝。
- 在中等约束强度下,违规回答率峰值超开放生成场景。
- 高能力模型生成的题目几乎全违规,且跨模型通用性强。
大语言模型的安全对齐通常在开放式生成中评估,此时模型可通过拒绝回答来规避风险。然而,在现实应用中,模型常被置于结构化决策任务中,如多选题(MCQ),此时拒绝选项被禁止或不可用。我们发现一种系统性失效模式:将有害请求改写为所有选项均不安全的强制多选题,可系统性绕过模型的拒绝行为,即使该模型在同等开放式提示下会一致拒绝。在14个私有和开源模型中,强制选择约束显著提升了违规响应率。值得注意的是,人工编写的多选题在中等约束强度下违规率呈倒U型分布,峰值出现在中间层级;而由高能力模型生成的多选题在各种约束下违规率接近饱和,且具有强跨模型迁移性。研究揭示当前安全评估严重低估了结构化任务中的风险,强调受限决策是未被充分探索的对齐失效关键界面。
原文摘要 · Abstract (English)
Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respond. In contrast, many real-world applications place LLMs in structured decision-making tasks, such as multiple-choice questions (MCQs), where abstention is discouraged or unavailable. We identify a systematic failure mode in this setting: reformulating harmful requests as forced-choice MCQs, where all options are unsafe, can systematically bypass refusal behavior, even in models that consistently reject equivalent open-ended prompts. Across 14 proprietary and open-source models, we show that forced-choice constraints sharply increase policy-violating responses. Notably, for human-authored MCQs, violation rates follow an inverted U-shaped trend with respect to structural constraint strength, peaking under intermediate task specifications, whereas MCQs generated by high-capability models yield near-saturation violation rates across constraints and exhibit strong cross-model transferability. Our findings reveal that current safety evaluations substantially underestimate risks in structured task settings and highlight constrained decision-making as a critical and underexplored surface for alignment failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。