arXiv:2510.08222cs.AI2025-10被引 1

从因果视角重构推理任务,提升大模型的逻辑能力。

Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens

  • 将推理视为选择机制,用逻辑概念筛选观察信息。
  • 在数独和迷宫任务上,性能提升超10%,参数减少8倍。
  • 适合研究大模型推理机制或想轻量化提升推理能力的人。

由于内在复杂性,推理任务长期被视为评估机器学习模型(尤其是大语言模型)能力的严格基准。尽管人类可轻松解决这些任务,现有模型即使经过大规模预训练和微调,仍无法可靠完成推理。本文从因果视角重新审视推理任务,旨在理解其在隐空间中的行为,并为应对挑战提供洞见。我们提出将推理任务建模为选择机制,其中高层逻辑概念作为选择算子作用于给定观测,如数学题中识别正确答案或数独中填入合适数值。该框架揭示两个关键特性:第一,隐空间复杂度超过观测空间,即使正确答案由输入完全决定;第二,对应逻辑思维的隐变量具有密集结构和强依赖关系。基于此,我们提出SR²框架,将估计的隐变量作为反馈引入选择机制,以促进隐表示间密集依赖的学习。该框架包含三个模块:反思表示学习、依赖自精炼与周期性中间对齐。实验表明,该方法在推理准确率上显著提升,在数独和迷宫任务上相比最新进展性能提升超10%,且参数量减少8倍。

原文摘要 · Abstract (English)

Due to their inherent complexity, reasoning tasks have long been regarded as rigorous benchmarks for assessing the capabilities of machine learning models, especially large language models (LLMs). Although humans can solve these tasks with ease, existing models, even after extensive pre-training and post-training at scale, still fail to perform reasoning reliably. In this paper, we revisit reasoning tasks from a causal perspective, seeking to understand their behavior in latent space and to offer insights for addressing their challenges. Specifically, we cast reasoning tasks as a selection mechanism, in which high-level logical concepts function as selection operators on the given observations, such as, identifying the correct answer in a math problem or filling the appropriate entry in Sudoku. We emphasize two key properties of this formulation that shed light on the difficulty of reasoning tasks. First, the latent space exceeds the observation space in complexity, even when the correct answer is fully determined by the observed input. Second, the latent variables, corresponding to logical thought, are densely structured and exhibit strong dependencies. Building on this formulation, we introduce a framework, called SR$^2$, that incorporates the estimated latent variables as feedback into the selection mechanism, thereby facilitating the learning of dense dependencies among latent representations. The framework consists of three key modules: reflective representation learning, dependency self-refinement, and periodic intermediate alignment. Experimentally, we show that our approach yields significant gains in reasoning accuracy, for example, attaining over 10$\%$ improvement in performance with 8$\times$ fewer parameters on the Sudoku and Maze tasks over the recent advances.

大模型推理因果建模逻辑推理轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。