用最优停止理论优化大模型自我修正次数,更省成本。
Optimal Stopping of Self-Refining Foundation Models
- 将模型自我修正过程建模为最优停止问题,按收益成本比决定迭代次数。
- 实验表明新策略在编码任务上比已有方法节省超30%计算资源。
- 适合关注推理效率与资源优化的研究者和工程实践者。
基础模型可通过外部反馈驱动的自修正过程提升输出质量。在此过程中,模型嵌入迭代循环:生成输出、接收验证器反馈,并通过上下文学习进行优化。本文提出一种新方法,将该过程形式化为最优停止问题,即根据预期改进与成本的权衡来决定修正迭代次数。我们推导出最优停止策略,并证明其可通过随机近似高效计算。为评估该方法,我们在基础模型代码生成基准上进行了实验。结果表明,所提出的停止策略在成本效率上显著优于先前工作中的策略。
原文摘要 · Abstract (English)
Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is embedded in an iterative loop where it generates outputs, receives feedback from verifiers, and refines its responses through in-context learning. Following a novel approach, we formalize this process as an optimal stopping problem where the number of refinement iterations is decided based on expected improvement relative to cost. We derive optimal stopping policies and show that they can be efficiently computed through stochastic approximation. To evaluate our approach experimentally, we apply it to a coding benchmark for foundation models. The empirical results show that our stopping policies are significantly more cost-efficient than stopping policies proposed in prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。