arXiv:2605.01772cs.ROcs.LG2026-05被引 2

通过动态生成子目标,让机器人更可靠地完成长时间任务。

Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation

论文配图:Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation
图 1 · 摘自论文原文
  • 用预测模型递归生成未来子目标,随任务进展动态调整。
  • 在仿真和真实机器人任务中,长程任务成功率显著提升。
  • 适合需要持续规划的机器人自主任务,如家庭服务、仓储搬运。

视觉-语言-动作(VLA)模型已成为具身智能的重要范式,使机器人能根据自然语言指令和当前视觉输入执行任务。然而,现有VLA模型在处理长时程任务时因误差累积而表现不佳。以往方法将任务分解为固定粒度的子任务,无法适应执行状态的复杂性变化,限制了其在长时程任务中的鲁棒性。为此,我们提出预测模型,可自适应地递归生成未来子目标。该模型在任务进行中持续调整,根据动态变化实时更新子目标,从而引导更可靠的规划路径。基于此思想,我们构建了层次化VLA模型Anticipation-VLA,利用预测模型生成可执行的子目标以指导VLA策略执行。我们通过微调统一多模态模型(UMM)实现高层子目标生成,并使用条件目标的VLA策略完成低层动作执行。在仿真与真实机器人任务中的实验验证了Anticipation-VLA的有效性,凸显了自适应递归子目标生成对鲁棒策略执行的重要性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for embodied intelligence, enabling robots to perform tasks based on natural language instructions and current visual input. However, existing VLA models struggle with long-horizon tasks due to compounding errors. Prior methods decompose tasks into subtasks of fixed granularity, which cannot adapt to the varying complexity of execution states, limiting their robustness in long-horizon tasks. To overcome this, we introduce Anticipation Model, which adaptively and recursively generates future subgoals. This model continuously adapts as the task unfolds, adjusting future subgoals in response to evolving dynamics, facilitating more reliable planning paths. Building on this concept, we propose Anticipation-VLA, a hierarchical VLA model that leverages the anticipation model to generate actionable subgoals that guide VLA policy execution. We implement Anticipation-VLA with finetuning a Unified Multimodal Model (UMM) for high-level subgoal generation and a goal-conditioned VLA policy for low-level action execution. Experiments in both simulated and real-world robotic tasks demonstrate the effectiveness of Anticipation-VLA, highlighting the importance of adaptive and recursive subgoal generation for robust policy execution.

具身智能子目标生成机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。