arXiv:2507.16473cs.AI2025-07

让大模型在隐空间中高效推理,不写步骤也能思考。

Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs

  • 用变分推断学习隐式思维动作,构建抽象决策框架。
  • 在逻辑推理和运动控制任务上达到顶尖性能,优于显式链式思考。
  • 适合追求高效推理的AI系统,尤其对资源受限场景友好。

大型语言模型通过显式的思维链提示展现出强大的推理能力,但生成每一步文本解释计算成本高且缓慢。为此,我们提出一种高效、隐式的推理框架,使模型在潜在空间中‘思考’,无需为每步生成显式文本。我们将这些潜在思维建模为层次强化学习框架中的时序抽象动作(即选项)。为有效学习多样化的选项作为潜在嵌入,我们引入变分马尔可夫选项评论家(VMOC),一种基于HiT-MDP框架的离策略算法,结合变分推断。为进一步证明该选项空间作为抽象推理空间的合理性,我们扩展了连续MDP同态理论,证明在简化后的抽象潜在空间中学习策略可保持原复杂问题解的最优性。最后,我们设计冷启动流程,利用监督微调数据将人类推理示范提炼至该潜在选项空间,为模型推理能力提供丰富初始化。大量实验表明,该方法在复杂逻辑推理基准和挑战性运动控制任务上表现优异,验证了其作为语言与控制领域抽象技能学习的原理性方法的有效性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable reasoning ability through explicit Chain-of-Thought (CoT) prompting, but generating these step-by-step textual explanations is computationally expensive and slow. To overcome this, we aim to develop a framework for efficient, implicit reasoning, where the model "thinks" in a latent space without generating explicit text for every step. We propose that these latent thoughts can be modeled as temporally-extended abstract actions, or options, within a hierarchical reinforcement learning framework. To effectively learn a diverse library of options as latent embeddings, we first introduce the Variational Markovian Option Critic (VMOC), an off-policy algorithm that uses variational inference within the HiT-MDP framework. To provide a rigorous foundation for using these options as an abstract reasoning space, we extend the theory of continuous MDP homomorphisms. This proves that learning a policy in the simplified, abstract latent space, for which VMOC is suited, preserves the optimality of the solution to the original, complex problem. Finally, we propose a cold-start procedure that leverages supervised fine-tuning (SFT) data to distill human reasoning demonstrations into this latent option space, providing a rich initialization for the model's reasoning capabilities. Extensive experiments demonstrate that our approach achieves strong performance on complex logical reasoning benchmarks and challenging locomotion tasks, validating our framework as a principled method for learning abstract skills for both language and control.

隐式推理强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。