arXiv:2602.04856cs.CL2026-02

模型拒绝生成假新闻,但其推理过程仍可能暗藏虚假信息。

CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation

  • 通过分析注意力头的响应模式,揭示推理中的潜在风险
  • 激活思考模式后生成风险显著上升,关键决策集中在少数中层
  • 为检测隐性虚假推理提供新方法,适合安全研究者参考

大型语言模型在生成标题甚至虚构新闻时,通常仅依据最终输出进行评估,普遍假设拒绝响应意味着整个推理过程安全。本文挑战这一假设,发现即使模型拒绝有害请求,其链式思维(CoT)过程中仍可能内部包含并传播不实叙事。为此,我们提出统一的安全分析框架,系统分解模型各层的CoT生成过程,并通过基于雅可比谱特征的度量方法,评估单个注意力头的作用。在此框架下,引入三个可解释指标:稳定性、几何性和能量,用以量化特定注意力头对欺骗性推理模式的响应与嵌入程度。在多个面向推理的LLM上进行的大量实验表明,启用思考模式后,生成风险显著上升,关键路由决策集中于少数连续的中深度层。通过精准定位导致分歧的注意力头,本工作质疑了‘拒绝即安全’的假设,为缓解潜在推理风险提供了新视角。

原文摘要 · Abstract (English)

From generating headlines to fabricating news, the Large Language Models (LLMs) are typically assessed by their final outputs, under the safety assumption that a refusal response signifies safe reasoning throughout the entire process. Challenging this assumption, our study reveals that during fake news generation, even when a model rejects a harmful request, its Chain-of-Thought (CoT) reasoning may still internally contain and propagate unsafe narratives. To analyze this phenomenon, we introduce a unified safety-analysis framework that systematically deconstructs CoT generation across model layers and evaluates the role of individual attention heads through Jacobian-based spectral metrics. Within this framework, we introduce three interpretable measures: stability, geometry, and energy to quantify how specific attention heads respond or embed deceptive reasoning patterns. Extensive experiments on multiple reasoning-oriented LLMs show that the generation risk rise significantly when the thinking mode is activated, where the critical routing decisions concentrated in only a few contiguous mid-depth layers. By precisely identifying the attention heads responsible for this divergence, our work challenges the assumption that refusal implies safety and provides a new understanding perspective for mitigating latent reasoning risks.

大模型安全推理机制虚假信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。