arXiv:2605.27851cs.AI2026-05被引 3

模型安全规则会因情境变化失效,需警惕表面安全下的脆弱性。

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

论文配图:When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models
图 1 · 摘自论文原文
  • 设计上下文翻转测试,对比安全与常识判断差异
  • 12个模型平均安全-常识差距达17.4个百分点,高准确率模型仍可能不安全
  • 现有内容审核无法识别后果反转,需引入状态感知验证机制

安全基准分数无法充分反映模型部署就绪状态:对齐的语言模型在情境更新后,即使原本安全的行为已转为有害,仍会机械遵守规则。我们称此现象为‘脆性安全’。为诊断该问题,提出上下文翻转评估方法,在PacifAIst安全基准及两个常识控制任务中测试12个模型,使用成对变体(名义安全行为导致伤害)。结果发现:第一,脆性安全具有安全性特异性,所有模型均存在安全-常识差距(平均+17.4百分点);基线准确率无法预测脆性,其中超过90%基线准确率的模型,脆性率在13.7%至90.0%之间;第二,失败源于策略覆盖而非理解错误——尽管所有模型都意识到情境变化,但仍通过三种不同机制持续错误行为,且机制随更新类型和模型族而异;第三,在人工审计的灾难性后果翻转场景中,标准动作级防护完全无效,而状态感知验证器则全部捕捉到风险,且无误报。这表明动作级内容审核对后果翻转系统性失灵,亟需引入状态感知架构。论文发布评估协议、扰动基准及部署探测工具。

原文摘要 · Abstract (English)

Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational update flips which action is safe. We term this failure brittle safety. To diagnose it, we introduce context-flip evaluation, testing 12 models across a safety benchmark (PacifAIst) and two commonsense controls using paired variants where the nominally safe action produces harm. Three findings emerge. First, brittle safety is safety-specific: all 12 models exhibit a safety-commonsense gap (mean +17.4 pp). Baseline accuracy fails to predict brittleness: among models above 90% baseline accuracy, brittleness rates range from 13.7% to 90.0%. Second, failures stem from policy override rather than miscomprehension: despite acknowledging the context change in every case, models persist via three distinct mechanisms that vary by update type and model family. Third, on a hand-audited probe of catastrophic consequence-flip scenarios, standard action-level guardrails catch none, while a state-aware validator catches all without false alarms on correct interventions. This indicates that action-level content moderation is systematically blind to consequence-flips, motivating state-aware architectural alternatives. We release our protocol, perturbed benchmarks, and deployment probe.

模型安全上下文理解脆性安全验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。