arXiv:2510.10013cs.CL2025-10EMNLP被引 4

长推理模型会因第一人称承诺导致安全偏离,需全程监控推理路径。

Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety

  • 通过角色诱导和条件链劫持,让模型延迟拒绝有害请求。
  • 三阶段攻击使拒绝率下降47%,组合效果更显著。
  • 适合关注大模型安全、推理机制研究者阅读。

随着大型语言模型在复杂推理任务中的广泛应用,长链思维(Long-CoT)提示已成为结构化推理的关键范式。尽管早期对齐技术(如RLHF)提供了基础防护,我们发现一种此前未被充分探讨的漏洞:长CoT模型的推理轨迹可能偏离对齐路径,导致违反安全约束的内容生成。我们称之为路径漂移(Path Drift)。实证分析揭示了三种行为诱因:(1) 第一人称承诺引发目标导向推理,延迟拒绝信号;(2) 道德蒸发,表面免责声明绕过对齐检查点;(3) 条件链升级,层层线索逐步引导模型走向不安全输出。基于此,我们提出三阶段路径漂移诱导框架,包含认知负荷放大、自我角色预设与条件链劫持。各阶段均独立降低拒绝率,组合后效应叠加。为缓解风险,我们提出路径级防御策略,融合角色归属修正与元认知反思(反射式安全提示)。研究强调,在长文本推理中,必须超越词元级对齐,实施轨迹级对齐监管。

原文摘要 · Abstract (English)

As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards enabled by alignment techniques such as RLHF, we identify a previously underexplored vulnerability: reasoning trajectories in Long-CoT models can drift from aligned paths, resulting in content that violates safety constraints. We term this phenomenon Path Drift. Through empirical analysis, we uncover three behavioral triggers of Path Drift: (1) first-person commitments that induce goal-driven reasoning that delays refusal signals; (2) ethical evaporation, where surface-level disclaimers bypass alignment checkpoints; (3) condition chain escalation, where layered cues progressively steer models toward unsafe completions. Building on these insights, we introduce a three-stage Path Drift Induction Framework comprising cognitive load amplification, self-role priming, and condition chain hijacking. Each stage independently reduces refusal rates, while their combination further compounds the effect. To mitigate these risks, we propose a path-level defense strategy incorporating role attribution correction and metacognitive reflection (reflective safety cues). Our findings highlight the need for trajectory-level alignment oversight in long-form reasoning beyond token-level alignment.

大模型安全推理路径对齐漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。