arXiv:2606.27268cs.ROcs.AI2026-06中稿 · ECCV

让机器人通过反思历史信息,实时优化操作决策。

E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation

论文配图:E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
图 1 · 摘自论文原文
  • 用带历史记忆的迭代反思机制,联合优化动作与推理。
  • 在仿真和真实场景中分别提升33.14%和26.62%性能。
  • 无需额外数据或训练,适配多种机器人与任务。

近期一些工作尝试研究具身任务的测试时扩展(TTS),但仍有两大挑战:(1)推理能有效提升策略性能,但其扩展机制少有研究;(2)具身任务具有长时序、顺序性特征,仅依赖当前观测进行动作扩展,缺乏对历史上下文的利用。为此,我们提出E-TTS,一种模块化、可插拔的具身测试时扩展框架,通过视觉-语言验证器实现历史感知的迭代优化,统一推理与动作扩展。为支持联合推理-动作扩展,E-TTS采用成对采样与评分机制。为更好利用历史信息,框架引入历史缓冲区,供推理与动作验证器评估候选动作。不同于传统开环TTS方法,E-TTS将反馈生成融入采样过程,形成闭环迭代优化机制,提升推理效率与环境适应性。各组件独立可组合,可根据任务需求灵活配置。我们在4个基准、6个环境、3种机器人形态及4个基础视觉-语言-动作模型上进行实验。结果表明,无需额外专家数据收集或重训练,E-TTS持续提升性能,在仿真中最高提升33.14%,真实世界提升26.62%。

原文摘要 · Abstract (English)

Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved: (1) reasoning can effectively improve the performance of the policy, but its scaling mechanism has seldom been studied; (2) historical information is essential, as embodied tasks are inherently long-horizon and sequential, making sole reliance on current observations for action scaling inadequate due to the lack of historical context utilization. To address these challenges, we introduce E-TTS, a modular and plug-and-play Embodied Test-Time Scaling framework that unifies reasoning and action scaling for robotic manipulation via history-aware iterative refinement with vision-language verifiers. To support joint reasoning-action scaling, E-TTS performs reasoning-action joint sampling and scoring in a pairwise manner. To better utilize historical information, E-TTS uses a history buffer to store historical context, which is then used by reasoning and action verifiers to evaluate the sampled candidates. Unlike conventional open-loop TTS methods, E-TTS introduces feedback generation into the sampling process to form a closed-loop iterative refinement mechanism, enhancing both inference efficiency and environmental adaptability. Each component functions as an independent and composable module, allowing flexible and adaptive configuration depending on task requirements. To evaluate the advantages of our framework, we conduct experiments across 4 different benchmarks, 6 environments, 3 embodiments, and 4 base vision-language-action models. The experimental results demonstrate that, without requiring additional expert data collection or retraining, E-TTS consistently improves performance, achieving up to a 33.14% increase in simulation and 26.62% in real-world scenarios.

机器人操控测试时扩展迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。