用历史经验提升推理效率,让大模型学会复用过往尝试。
Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework
- 通过'贴纸'机制提取关键条件,指导多轮推理迭代
- 在AIME-24/25和OlymMATH上超越自洽与强化学习基线
- 适合追求高推理效率的复杂任务研究者
大型推理模型(LRMs)在复杂推理任务中表现优异,进一步提升可通过增加推理时计算预算实现。然而,现有测试时扩展方法主要依赖冗余采样,忽视历史经验利用,限制了计算效率。为此,我们提出Sticker-TTS,一种新型测试时扩展框架,通过三个协作的大型推理模型,基于历史尝试迭代探索并优化解决方案。框架核心是提炼出的、被称为'贴纸'的关键条件,驱动跨多轮推理中关键信息的提取、精炼与复用。为进一步提升效率与性能,引入两阶段优化策略,结合模仿学习与自我改进,实现渐进式优化。在三个具有挑战性的数学推理基准(包括AIME-24、AIME-25和OlymMATH)上的大量评估表明,Sticker-TTS在相当的推理预算下持续优于强基线,包括自洽性和先进强化学习方法。结果凸显了贴纸引导的历史经验利用的有效性。代码与数据可在https://github.com/RUCAIBox/Sticker-TTS获取。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) have exhibited strong performance on complex reasoning tasks, with further gains achievable through increased computational budgets at inference. However, current test-time scaling methods predominantly rely on redundant sampling, ignoring the historical experience utilization, thereby limiting computational efficiency. To overcome this limitation, we propose Sticker-TTS, a novel test-time scaling framework that coordinates three collaborative LRMs to iteratively explore and refine solutions guided by historical attempts. At the core of our framework are distilled key conditions-termed stickers-which drive the extraction, refinement, and reuse of critical information across multiple rounds of reasoning. To further enhance the efficiency and performance of our framework, we introduce a two-stage optimization strategy that combines imitation learning with self-improvement, enabling progressive refinement. Extensive evaluations on three challenging mathematical reasoning benchmarks, including AIME-24, AIME-25, and OlymMATH, demonstrate that Sticker-TTS consistently surpasses strong baselines, including self-consistency and advanced reinforcement learning approaches, under comparable inference budgets. These results highlight the effectiveness of sticker-guided historical experience utilization. Our code and data are available at https://github.com/RUCAIBox/Sticker-TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。