arXiv:2510.04320cs.CLcs.LG2025-10被引 5

让大模型学会看行为后果,而非只看表面话术。

Read the Scene, Not the Script: Outcome-Aware Safety for LLMs

  • 构建新基准CB-Bench,区分语义风险与实际后果
  • 用新数据集训练后,抗越狱能力提升,误拒减少
  • 适合研究模型安全对齐与真实场景泛化的人

当前安全对齐的大语言模型仍存在两大缺陷:易被越狱,或对无害输入过度拒绝。我们发现根源在于模型对行为与结果之间的关联推理薄弱,过度依赖语义或风格等表面线索。为此,我们提出“后果盲视”(Consequence-blindness)概念,并构建涵盖四种风险场景的基准CB-Bench,可评估语义风险与实际后果匹配与否的情况。主流模型在多种条件下均表现不佳,说明该问题普遍存在。我们进一步推出用于安全对齐的后果推理数据集CS-Chain-4k,基于该数据集微调的模型显著提升了对语义伪装越狱的防御能力,同时降低对无害输入的误拒率,且在其他基准上保持良好性能。这揭示了现有对齐方法的局限,确立后果感知推理为核心目标,并提供更实用、可复现的评估路径。

原文摘要 · Abstract (English)

Safety-aligned Large Language Models (LLMs) still show two dominant failure modes: they are easily jailbroken, or they over-refuse harmless inputs that contain sensitive surface signals. We trace both to a common cause: current models reason weakly about links between actions and outcomes and over-rely on surface-form signals, lexical or stylistic cues that do not encode consequences. We define this failure mode as Consequence-blindness. To study consequence-blindness, we build a benchmark named CB-Bench covering four risk scenarios that vary whether semantic risk aligns with outcome risk, enabling evaluation under both matched and mismatched conditions which are often ignored by existing safety benchmarks. Mainstream models consistently fail to separate these risks and exhibit consequence-blindness, indicating that consequence-blindness is widespread and systematic. To mitigate consequence-blindness, we introduce CS-Chain-4k, a consequence-reasoning dataset for safety alignment. Models fine-tuned on CS-Chain-4k show clear gains against semantic-camouflage jailbreaks and reduce over-refusal on harmless inputs, while maintaining utility and generalization on other benchmarks. These results clarify the limits of current alignment, establish consequence-aware reasoning as a core alignment goal and provide a more practical and reproducible evaluation path.

模型安全因果推理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。