arXiv:2608.05409cs.CL2026-08

语法形式影响模型拒绝能力,安全对齐存在隐藏漏洞

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

论文配图:Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
图 1 · 摘自论文原文
  • 发现非命令句式会削弱模型安全拒绝能力
  • 16个模型(最大70B参数)均存在此缺陷
  • 适合关注AI安全与对齐机制的研究者

大型语言模型通常通过后训练对齐安全策略,但存在多种复杂越狱方法可绕过防护。例如,Andriushchenko等人(2025)发现将时态从现在时改为过去时即可诱导有害响应。本文揭示了更普遍的非命令句式导致的安全失效问题。我们在16个模型(最大70B参数)上通过行为评估验证了该现象。因果中介分析显示,拒绝行为部分依赖于上游句法特征。通过操控这些纯句法特征,可主动触发或抑制拒绝。进一步追踪发现,这是由开源模型后训练数据的语言偏差所致,增加句法多样性可缓解该问题。结果表明,当前对齐方法引入混淆因素,阻碍拒绝决策的纯粹语义基础。

原文摘要 · Abstract (English)

Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.

安全对齐句法敏感模型漏洞后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。