arXiv:2510.16340cs.CLcs.AI2025-10被引 1

评估大模型在后训练后的思维意识与推理一致性。

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

  • 设计三类能力测试,考察模型对自身策略的认知
  • 强化学习训练模型更懂自己,但推理过程常与答案不一致
  • 适合关注大模型可解释性与可信推理的研究者

近期的后训练技术使大语言模型通过生成额外规划令牌,在处理复杂逻辑任务时表现增强。这引发一个根本问题:这些模型是否意识到自己“学到了什么”和“在思考什么”?为此,我们定义了三个核心能力:(1)对所学隐含策略的自我认知,(2)在不同领域间策略泛化能力,(3)内部推理轨迹与最终输出的一致性。我们在多个任务上实证评估这些能力,每个任务需学习不同策略。对比了经监督微调(SFT)、直接策略优化(DPO)和组相对策略优化(GRPO)后训练的模型。结果表明,强化学习训练的模型不仅比SFT模型更清楚自身行为、在新结构相似任务中泛化更强,且普遍表现出推理轨迹与输出之间一致性较弱,这一现象在GRPO训练模型中尤为明显。

原文摘要 · Abstract (English)

Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensive tasks through the generation of supplementary planning tokens. This development raises a fundamental question: Are these models aware of what they "learn" and "think"? To address this, we define three core competencies: (1) awareness of learned latent policies, (2) generalization of these policies across domains, and (3) alignment between internal reasoning traces and final outputs. We empirically evaluate these abilities on several tasks, each designed to require learning a distinct policy. Furthermore, we contrast the profiles of models post-trained via Supervised Fine-Tuning (SFT), Direct Policy Optimization (DPO), and Group Relative Policy Optimization (GRPO). Our findings indicate that RL-trained models not only demonstrate greater awareness of their learned behaviors and stronger generalizability to novel, structurally similar tasks than SFT models but also often exhibit weak alignment between their reasoning traces and final outputs, an effect most pronounced in GRPO-trained models.

大模型推理可解释性强化学习后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。