arXiv:2412.17256cs.AIcs.CL2024-12ICLR被引 30

提出自适应调节探索与利用的框架,提升模型自我改进效率

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

  • 动态监测模型生成多样性与奖励区分能力
  • 迭代中探索能力下降导致性能瓶颈,奖励效果减弱
  • 适用于需要持续优化的复杂推理任务,如数学与编程

在缺乏大量人工标注数据的复杂推理任务中,基于自身输出进行训练的自我改进方法已成为提升模型性能的主要途径。然而,这类迭代式自我改进机制的关键因素仍不明确,例如何时有效、当前迭代中的瓶颈为何。本文以数学推理为例,通过定量分析发现:模型的探索能力在迭代过程中迅速衰退,外部奖励区分优质与劣质候选的能力也随之下降。针对此问题,我们提出B-STaR框架,可自主调整配置以平衡探索与利用,根据当前策略模型和可用奖励优化自我改进效果。在数学推理、编码和常识推理任务上的实验表明,B-STaR不仅提升了整个训练过程中的探索能力,还实现了更优的探索-利用平衡,显著提升性能。

原文摘要 · Abstract (English)

In the absence of extensive human-annotated data for complex reasoning tasks, self-improvement -- where models are trained on their own outputs -- has emerged as a primary method for enhancing performance. However, the critical factors underlying the mechanism of these iterative self-improving methods remain poorly understood, such as under what conditions self-improvement is effective, and what are the bottlenecks in the current iterations. In this work, we identify and propose methods to monitor two pivotal factors in this iterative process: (1) the model's ability to generate sufficiently diverse responses (exploration); and (2) the effectiveness of external rewards in distinguishing high-quality candidates from lower-quality ones (exploitation). Using mathematical reasoning as a case study, we begin with a quantitative analysis to track the dynamics of exploration and exploitation, discovering that a model's exploratory capabilities rapidly deteriorate over iterations, and the effectiveness of exploiting external rewards diminishes as well. Motivated by these findings, we introduce B-STaR, a Self-Taught Reasoning framework that autonomously adjusts configurations across iterations to Balance exploration and exploitation, thereby optimizing the self-improving effectiveness based on the current policy model and available rewards. Our experiments on mathematical reasoning, coding, and commonsense reasoning demonstrate that B-STaR not only enhances the model's exploratory capabilities throughout training but also achieves a more effective balance between exploration and exploitation, leading to superior performance.

自我改进推理增强探索利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。