arXiv:2512.05962cs.LGcs.AI2025-12被引 3

通过过滤错误答案,提升大模型推理多样性与精度。

Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity

  • 用过滤法构建目标分布,避免强化学习导致的模式集中
  • 在定理证明任务上实现最高覆盖率和最优精度平衡
  • 适合需要多样化推理路径的复杂任务研究者

强化学习已成为调优大语言模型进行推理任务的标准方法。然而,越来越多证据表明,此类训练方式会导致模型多样性显著下降。我们提出,这是由于强化学习隐式优化了反向KL散度(零强迫),使模型过度集中在高概率区域而忽略其他正确答案。本文从显式的目标分布出发,通过过滤错误答案并保留正确答案间的相对概率来构建目标。基于预训练模型,我们采用α-散度族近似该分布,统一已有方法,并可通过插值控制精度与多样性的权衡。在精益定理证明基准测试中,该方法在覆盖-精度帕累托前沿上达到当前最优,且在覆盖率方面超越所有已有方法。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has become the de facto standard for tuning LLMs to solve tasks involving reasoning. However, growing evidence shows that models trained in such way often suffer from a significant loss in diversity. We argue that this arises because RL implicitly optimizes the "mode-seeking" or "zero-forcing" Reverse KL to a target distribution causing the model to concentrate mass on certain high-probability regions of the target while neglecting others. In this work, we instead begin from an explicit target distribution, obtained by filtering out incorrect answers while preserving the relative probabilities of correct ones. Starting from a pre-trained LLM, we approximate this target distribution using the $α$-divergence family, which unifies prior approaches and enables direct control of the precision-diversity trade-off by interpolating between mode-seeking and mass-covering divergences. On a Lean theorem-proving benchmark, our method achieves state-of-the-art performance along the coverage-precision Pareto frontier, outperforming all prior methods on the coverage axis.

推理增强模型多样性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。