让大模型先看答案,反推高质量推理路径。
RAVR: Reference-Answer-guided Variational Reasoning for Large Language Models
- 以答案为条件引导推理,提升路径质量
- 数学与通用任务均显著优于基线
- 适合提升模型逻辑推理能力的场景
强化学习可提升大语言模型的推理能力,但依赖于模型已能以非可忽略概率生成高价值推理路径。对于超出当前能力的任务,此类路径难以采样,学习可能固化低效推理模式。受认知科学启发:解释‘为何是这个答案’比‘答案是什么’更易,因后者避免开放式探索,转而系统重构问题与答案间的推理链条。我们证明,大模型可借助答案推导出高质量推理路径;在答案条件下采样,可严格提高所采路径的期望效用,使原本不可解的问题变为可学。基于此,我们提出 RAVR(参考答案引导的变分推理)框架,采用答案条件化推理作为仅问句推理的变分替代。在通用及数学领域实验中,结果持续优于强基线。进一步分析显示,RAVR减少犹豫、强化结论整合,并促进特定问题策略的形成。
原文摘要 · Abstract (English)
Reinforcement learning (RL) can refine the reasoning abilities of large language models (LLMs), but critically depends on a key prerequisite: the LLM can already generate high-utility reasoning paths with non-negligible probability. For tasks beyond the LLM's current competence, such reasoning path can be hard to sample, and learning risks reinforcing familiar but suboptimal reasoning. We are motivated by the insight from cognitive science that Why is this the answer is often an easier question than What is the answer, as it avoids the heavy cognitive load of open-ended exploration, opting instead for explanatory reconstruction-systematically retracing the reasoning that links a question to its answer. We show that LLMs can similarly leverage answers to derive high-quality reasoning paths. We formalize this phenomenon and prove that conditioning on answer provably increases the expected utility of sampled reasoning paths, thereby transforming intractable problems into learnable ones. Building on this insight, we introduce RAVR (Reference-Answer-guided Variational Reasoning), an end-to-end framework that uses answer-conditioned reasoning as a variational surrogate for question-only reasoning. Experiments in both general and math domains demonstrate consistent improvements over strong baselines. We further analyze the reasoning behavior and find that RAVR reduces hesitation, strengthens conclusion consolidation, and promotes problem-specific strategies in reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。