arXiv:2512.12576cs.CLcs.AI2025-12中稿 · ICML被引 1

让大模型推理更连贯高效,通过联合优化思维路径与答案

Coupled Variational Reinforcement Learning for Language Model General Reasoning

  • 用混合采样策略将推理路径与答案信息耦合,提升探索效率
  • 在数学和通用推理任务上比基线模型提升12.4%,超越现有无验证器方法2.3%
  • 适合需要强逻辑推理能力的场景,如数学题求解、复杂问题分析

尽管强化学习在语言模型推理中取得显著进展,但受限于可验证奖励的需求。近期无验证器的强化学习方法利用大模型生成参考答案的概率作为奖励信号,克服了这一限制。然而,这些方法通常仅基于问题采样推理路径,导致推理路径与最终答案脱节,影响探索效率和逻辑一致性。本文提出耦合变分强化学习(CoVRL),通过混合采样策略将先验与后验分布耦合,构建并优化融合两者的复合分布。该方法在保持强思维-答案一致性的同时实现高效探索。在数学和通用推理基准上的大量实验表明,CoVRL相较于基线模型性能提升12.4%,并在无验证器方法中再提升2.3%,为增强语言模型的通用推理能力提供了理论严谨的框架。

原文摘要 · Abstract (English)

While reinforcement learning has achieved impressive progress in language model reasoning, it is constrained by the requirement for verifiable rewards. Recent verifier-free RL methods address this limitation by utilizing the probabilities that LLMs generate reference answers as reward signals. However, these approaches typically sample reasoning traces conditioned only on the question. This design decouples reasoning-trace sampling from answer information, leading to inefficient exploration and incoherence between traces and final answers. In this paper, we propose \textit{\b{Co}upled \b{V}ariational \b{R}einforcement \b{L}earning} (CoVRL), which bridges variational inference and reinforcement learning by coupling prior and posterior distributions through a hybrid sampling strategy. By constructing and optimizing a composite distribution that integrates these two distributions, CoVRL enables efficient exploration while preserving strong thought-answer coherence. Extensive experiments on mathematical and general reasoning benchmarks show that CoVRL improves performance by 12.4\% over the base model and achieves an additional 2.3\% improvement over state-of-the-art verifier-free RL baselines, providing a principled framework for enhancing the general reasoning capabilities of language models.

强化学习逻辑推理大模型变分推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。