arXiv:2608.16885cs.RO2026-08

让机器人在执行长任务时能动态分配计算资源,提升决策准确性和成功率。

$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

论文配图:$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
图 1 · 摘自论文原文
  • 通过世界模型引导,在测试阶段动态增加复杂决策的计算量。
  • 在真实数据上训练40,115小时,支持跨机器人形态执行任务。
  • 适合需要长期规划与高可靠性的机器人控制场景。

长周期机器人操作需要可靠执行单个技能并合理编排这些技能完成复杂任务。现有分层视觉-语言-动作(VLA)模型通常仅用一次前向传播做决策,无法为关键或困难选择增加计算。本文提出$τ_0$-VLA,一种分层机器人基础模型,将高层子任务生成建模为可扩展计算的推理问题,通过世界模型引导的测试时计算实现动态资源分配。每个推理步骤中,高层策略利用执行记忆生成子任务,并在必要时搜索多个备选方案后确定输出;低层策略则在多个机器人本体上执行生成的子任务。模型基于40,115小时异构真实世界数据进行多模态联合训练。在域内及分布外设置下,增加测试时计算显著提升了下一子任务预测准确率,且这些收益转化为更高闭环成功率,显著提升长周期机器人操作性能。

原文摘要 · Abstract (English)

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $τ_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

机器人分层决策测试时计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。