arXiv:2508.19827cs.AIcs.CL2025-08EMNLP被引 17

探究思维链在软性推理中的真实作用,发现其效果与模型类型密切相关。

Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?

  • 对比指令微调、推理和推理蒸馏模型的思维链行为差异
  • 发现思维链影响与模型实际推理不一致,存在后验解释偏差
  • 适合关注大模型推理可信性的研究人员参考

近期研究表明,思维链(CoT)在分析性与常识推理等软性推理任务中提升有限,且可能无法忠实反映模型的真实推理过程。本文研究了指令微调、推理型及推理蒸馏模型在软性推理任务中思维链的作用机制与忠实度。结果表明,不同模型对思维链的依赖程度存在显著差异,且思维链的影响与其真实性并非总呈正相关,揭示了思维链在实际应用中的复杂动态。

原文摘要 · Abstract (English)

Recent work has demonstrated that Chain-of-Thought (CoT) often yields limited gains for soft-reasoning problems such as analytical and commonsense reasoning. CoT can also be unfaithful to a model's actual reasoning. We investigate the dynamics and faithfulness of CoT in soft-reasoning tasks across instruction-tuned, reasoning and reasoning-distilled models. Our findings reveal differences in how these models rely on CoT, and show that CoT influence and faithfulness are not always aligned.

思维链推理分析模型可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。