arXiv:2510.22362cs.LGcs.CL2025-10中稿 · NeurIPS被引 1

提出概念行走框架,检测大模型推理是否真实可靠。

Mapping Faithful Reasoning in Language Models

  • 在激活空间中追踪模型对概念的内部立场变化
  • 简单任务中推理被快速忽略,复杂任务中激活持续改变
  • 帮助识别装饰性推理,适合模型可信性研究者使用

链式思维(CoT)本应提升语言模型推理的可解释性,但已有研究表明其未必忠实反映内部计算过程。这导致实践者可能误将装饰性推理当作真实推理。本文提出概念行走(Concept Walk)框架,通过对比数据学习的概念方向,在激活空间中追踪推理过程中模型内部立场的演变。以 Qwen 3-4B 模型在安全领域为例,发现:在“简单”任务中,扰动后的 CoT 被迅速忽略,表明为装饰性推理;而在“困难”任务中,扰动引发内部激活的持续变化,符合真实推理特征。该方法从概念层面揭示内部动态,有助于判断推理是否可信。

原文摘要 · Abstract (English)

Chain-of-thought (CoT) traces promise transparency for reasoning language models, but prior work shows they are not always faithful reflections of internal computation. This raises challenges for oversight: practitioners may misinterpret decorative reasoning as genuine. We introduce Concept Walk, a general framework for tracing how a model's internal stance evolves with respect to a concept direction during reasoning. Unlike surface text, Concept Walk operates in activation space, projecting each reasoning step onto the concept direction learned from contrastive data. This allows us to observe whether reasoning traces shape outcomes or are discarded. As a case study, we apply Concept Walk to the domain of Safety using Qwen 3-4B. We find that in 'easy' cases, perturbed CoTs are quickly ignored, indicating decorative reasoning, whereas in 'hard' cases, perturbations induce sustained shifts in internal activations, consistent with faithful reasoning. The contribution is methodological: Concept Walk provides a lens to re-examine faithfulness through concept-specific internal dynamics, helping identify when reasoning traces can be trusted and when they risk misleading practitioners.

推理可信性激活空间模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。