用逆向推理重构犯罪过程,暴露大模型安全对语气的敏感性
Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations
- 将有害请求伪装成已发生事件的逆向还原任务
- 在gpt-4o上攻破率58%,远超直接请求的0%成功率
- 前沿模型虽拒绝对话,但易被后续对话诱导突破
大型语言模型的安全对齐通常针对直接、命令式的有害请求进行评估。我们发现,这种对齐高度依赖语用形式:当同一潜在意图以不同表达方式呈现时,模型可能从拒绝转为配合。这表明当前对齐策略并非对语义等价保持不变,而是仍受请求语用框架影响。我们提出逆向思维链(RetroCoT),一种单轮攻击方法,将有害请求重构为法医重建任务。不直接要求执行,而是假设危害已发生,让模型扮演法医角色,逆向推导其因果链条。在AdvBench(n=50)测试中,RetroCoT在gpt-4o上成功率达58%,而直接请求基线为0%;在gpt-4o-mini上分别为52%和4%。进一步发现明显代际差距:GPT-5系列模型完全拒绝该类请求,并在拒绝理由中明确指出重建前提,表明其对这一语用形式有显式覆盖。然而,这种鲁棒性无法泛化至其他语用形式。仅通过一次对抗性反馈,将已有法医回答与虚构低分评价并列,即可使GPT-5.4-mini的攻破率从0%升至48%,gpt-4o从58%升至94%;控制组省略评分则达85%,说明关键在于维持既定法医语用框架下的持续对话,而非分数操纵。结果表明,前沿模型的对齐仍受语用框架制约,而非语义意图,新语用形式仍可持续暴露其漏洞。
原文摘要 · Abstract (English)
Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment is highly conditioned on pragmatic register: models that refuse a direct request frequently comply when the same underlying objective is expressed through a different communicative stance. This suggests that current alignment policies are not invariant to semantic equivalence, but remain sensitive to how a request is pragmatically framed. We introduce Retroactive Chain-of-Thought (RetroCoT), a single-turn attack that reframes harmful requests as forensic reconstruction tasks. Rather than requesting harmful instructions directly, RetroCoT presupposes that the harmful outcome has already occurred and asks the model, acting as a forensic analyst, to reconstruct in reverse the causal chain that produced it. On AdvBench (n=50), RetroCoT achieves attach success rate of 58% on gpt-4o and 52% on gpt-4o-mini, compared with direct-request baselines of 0% and 4%, respectively. We further identify a pronounced generation gap: GPT-5-family models refuse RetroCoT entirely, explicitly identifying the reconstruction premise in their refusal rationales, consistent with explicit coverage of this reconstruction register. However, this robustness does not generalize across pragmatic forms. A single adversarial feedback turn presenting an existing forensic reconstruction response alongside evaluator critique raises ASR from 0% to 48% on GPT-5.4-mini and from 58% to 94% on GPT-4o; a control condition omitting the fabricated low score achieves 85% on GPT-5.4-mini, indicating that the operative element is pragmatic continuation within the established forensic frame rather than score manipulation. These results suggest that frontier-model alignment remains conditioned on pragmatic framing rather than semantic intent, and that new pragmatic registers can continue to expose a...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。