arXiv:2603.12397cs.CL2026-03被引 1

推理过程本身会因果影响模型行为,即使答案相同。

Not Just the Destination, But the Journey: Reasoning Traces Causally Shape Generalization Behaviors

  • 设计对照实验,固定有害答案,改变推理路径。
  • 不同推理类型导致不同行为模式,即便最终答案一致。
  • 无需答案监督,仅训练推理就能改变模型行为,适合对齐研究者看。

链式思维(CoT)常被视为大模型决策的窗口,但近期研究认为其可能仅为事后合理化。这引发关键问题:推理过程是否在不依赖最终答案的情况下,独立地因果影响模型泛化能力?为隔离推理的因果效应,我们设计控制实验,保持有害答案不变,而改变推理路径。构建包含恶意(Evil)、误导性(Misleading)和顺从性(Submissive)推理的数据集。在0.6B至14B参数范围内,采用问题-思考-答案(QTA)、问题-思考(QT)、仅思考(T-only)等范式训练模型,并在有思考与无思考模式下评估。结果表明:(1)CoT训练比标准微调更易放大有害泛化;(2)不同推理类型诱发与语义一致的行为模式,即使最终答案相同;(3)仅用推理进行训练(如QT或T-only)即可改变行为,证明推理携带独立信号;(4)这些影响在无推理生成答案时仍持续存在,表明其已被深度内化。研究证实推理内容具有因果效力,挑战仅监督输出的对齐策略。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) is often viewed as a window into LLM decision-making, yet recent work suggests it may function merely as post-hoc rationalization. This raises a critical alignment question: Does the reasoning trace causally shape model generalization independent of the final answer? To isolate reasoning's causal effect, we design a controlled experiment holding final harmful answers constant while varying reasoning paths. We construct datasets with \textit{Evil} reasoning embracing malice, \textit{Misleading} reasoning rationalizing harm, and \textit{Submissive} reasoning yielding to pressure. We train models (0.6B--14B parameters) under multiple paradigms, including question-thinking-answer (QTA), question-thinking (QT), and thinking-only (T-only), and evaluate them in both think and no-think modes. We find that: (1) CoT training could amplify harmful generalization more than standard fine-tuning; (2) distinct reasoning types induce distinct behavioral patterns aligned with their semantics, despite identical final answers; (3) training on reasoning without answer supervision (QT or T-only) is sufficient to alter behavior, proving reasoning carries an independent signal; and (4) these effects persist even when generating answers without reasoning, indicating deep internalization. Our findings demonstrate that reasoning content is causally potent, challenging alignment strategies that supervise only outputs.

大模型对齐链式思维行为泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。