用细粒度价值信号优化大模型推理,无需人工标注就能提升思维链表现。
Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values
- 通过每一步的值信号替代偏好标签,用均方误差直接优化模型。
- 在数学与常识推理任务中,用更少训练步数超越现有离线偏好优化方法。
- 适合缺乏人工标注数据的场景,尤其适用于复杂推理能力提升。
我们提出直接价值优化(DVO),一种用于提升大语言模型在复杂推理任务中表现的强化学习框架。不同于依赖偏好标签的传统方法,DVO利用推理过程各步骤的价值信号,通过均方误差损失进行优化。其核心优势在于细粒度监督,避免了耗时的人工标注。DVO中的目标值可通过蒙特卡洛树搜索或结果值模型估计。我们在数学与常识推理任务上的实证分析表明,即使训练步数更少,DVO也持续优于现有的离线偏好优化技术。这些发现凸显了价值信号在推动推理能力发展中的重要性,并证明了DVO在缺乏明确人类偏好信息时的优越性。
原文摘要 · Abstract (English)
We introduce Direct Value Optimization (DVO), an innovative reinforcement learning framework for enhancing large language models in complex reasoning tasks. Unlike traditional methods relying on preference labels, DVO utilizes value signals at individual reasoning steps, optimizing models via a mean squared error loss. The key benefit of DVO lies in its fine-grained supervision, circumventing the need for labor-intensive human annotations. Target values within the DVO are estimated using either Monte Carlo Tree Search or an outcome value model. Our empirical analysis on both mathematical and commonsense reasoning tasks shows that DVO consistently outperforms existing offline preference optimization techniques, even with fewer training steps. These findings underscore the importance of value signals in advancing reasoning capabilities and highlight DVO as a superior methodology under scenarios lacking explicit human preference information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。