arXiv:2412.10345cs.ROcs.AI2024-12ICLR被引 311

用视觉轨迹提示提升机器人模型对时空动态的感知能力

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

论文配图:TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
图 1 · 摘自论文原文
  • 通过可视化状态-动作轨迹增强VLA模型的时空感知
  • 在仿真和真实机器人上分别提升10%和3.5倍性能
  • 适合需要强泛化能力的通用机器人策略研究

尽管基于大规模机器人数据预训练的视觉-语言-动作(VLA)模型为机器人学习提供了有前景的通用策略,但在交互式机器人中的时空动态建模仍存在困难,影响复杂任务(如操作)的表现。本文提出视觉轨迹提示方法,通过可视化编码状态-动作轨迹来增强VLA模型对时空动态的感知。我们基于自收集的15万条机器人操作轨迹,微调OpenVLA构建了新模型TraceVLA。在SimplerEnv的137种配置及物理机器人WidowX上的4项任务中评估显示,TraceVLA在仿真环境中比OpenVLA提升10%,在真实机器人任务中提升3.5倍,且在多种机器人形态与场景下表现出强大泛化能力。为进一步验证方法的有效性与通用性,我们还基于4B Phi-3-Vision构建了一个紧凑型VLA模型,在Open-X-Embodiment上预训练并微调于本数据集,其性能媲美7B OpenVLA,同时显著提升推理效率。

原文摘要 · Abstract (English)

Although large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive robotics, making them less effective in handling complex tasks, such as manipulation. In this work, we introduce visual trace prompting, a simple yet effective approach to facilitate VLA models' spatial-temporal awareness for action prediction by encoding state-action trajectories visually. We develop a new TraceVLA model by finetuning OpenVLA on our own collected dataset of 150K robot manipulation trajectories using visual trace prompting. Evaluations of TraceVLA across 137 configurations in SimplerEnv and 4 tasks on a physical WidowX robot demonstrate state-of-the-art performance, outperforming OpenVLA by 10% on SimplerEnv and 3.5x on real-robot tasks and exhibiting robust generalization across diverse embodiments and scenarios. To further validate the effectiveness and generality of our method, we present a compact VLA model based on 4B Phi-3-Vision, pretrained on the Open-X-Embodiment and finetuned on our dataset, rivals the 7B OpenVLA baseline while significantly improving inference efficiency.

机器人策略视觉提示时空感知泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。