通过三维扩展提升推理模型测试时表现,突破上下文长度限制
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
- 引入上下文、批量、迭代三维度测试时扩展框架
- 在IOI、IMO等难题上显著提升推理准确率,支持人类反馈优化
- 适用于复杂任务与具身学习场景,拓展应用边界
推理增强型强化学习近期揭示了新的缩放效应:测试时缩放。如R1和o1等思考型模型在测试时随着推理上下文长度增加,准确率得以提升。然而,相较于训练时缩放,测试时缩放受基础模型上下文长度限制,远低于训练时消耗的令牌数量。本文从缩放效应角度重新审视测试时增强技术,提出统一的多维测试时缩放框架,拓展测试时推理能力。除传统上下文长度缩放外,还引入两个新维度:批量缩放(并行采样提升准确率)与回合缩放(迭代自修正改善推理质量)。基于此,提出三维测试时缩放,整合上下文、批量与回合缩放。实验表明:(1) 每个维度均呈现测试时缩放效应,但容量有限;(2) 三者结合显著提升复杂测试基准(如IOI、IMO、CPHO)的推理性能,并进一步受益于人类偏好反馈;(3) 人机协同框架自然延伸至更开放领域——具身学习,实现类人控制行为设计。
原文摘要 · Abstract (English)
Reasoning reinforcement learning (RL) has recently revealed a new scaling effect: test-time scaling. Thinking models such as R1 and o1 improve their reasoning accuracy at test time as the length of the reasoning context increases. However, compared with training-time scaling, test-time scaling is fundamentally limited by the limited context length of base models, which remains orders of magnitude smaller than the amount of tokens consumed during training. We revisit test-time enhancement techniques through the lens of scaling effect and introduce a unified framework of multi-dimensional test-time scaling to extend the capacity of test-time reasoning. Beyond conventional context-length scaling, we consider two additional dimensions: batch scaling, where accuracy improves with parallel sampling, and turn scaling, where iterative self-refinement enhances reasoning quality. Building on this perspective, we propose 3D test-time scaling, which integrates context, batch, and turn scaling. We show that: (1) each dimension demonstrates a test-time scaling effect, but with a bounded capacity; (2) combining all three dimensions substantially improves the reasoning performance of challenging testbeds, including IOI, IMO, and CPHO, and further benefits from human preference feedback; and (3) the human-in-the-loop framework naturally extends to a more open-ended domain, i.e., embodied learning, which enables the design of humanoid control behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。