用多种引导信号提升机器人通用策略在复杂任务中的表现
OmniGuide: Universal Guidance Fields for Enhancing Generalist Robot Policies
- 将各类引导信息转化为三维空间中的可微能量函数
- 在仿真与真实环境中显著提升成功率和安全性
- 适合作为增强现有机器人策略的通用插件框架
视觉-语言-动作(VLA)模型在众多简单任务中表现出色,但在需要复杂空间或语义理解、杂乱环境操作或精确控制的复杂任务上性能有限。本文提出OMNIGUIDE,一种灵活框架,通过利用3D基础模型、语义推理视觉语言模型、人体姿态模型等任意引导源,提升VLA在复杂任务上的表现。我们展示如何将多种引导信息自然表示为带有任务特定吸引子和排斥子的可微能量函数,影响VLA动作采样。该方法使互补优势的引导源协同优化模型性能。在仿真与真实世界环境中的大量实验表明,OMNIGUIDE显著提升当前顶尖通用策略(如π_{0.5}、GR00T N1.6)的成功率与安全率。关键的是,该统一框架在性能上匹配甚至超越了针对特定引导源设计的先前方法。
原文摘要 · Abstract (English)
Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on more complex tasks, such as those requiring complex spatial or semantic understanding, manipulation in clutter, or precise manipulation. We propose OMNIGUIDE, a flexible framework that improves VLA performance on such tasks by leveraging arbitrary sources of guidance, such as 3D foundation models, semantic-reasoning VLMs, and human pose models. We show how many kinds of guidance can be naturally expressed as differentiable energy functions with task-specific attractors and repellers located in 3D space, that influence the sampling of VLA actions. In this way, OMNIGUIDE enables guidance sources with complementary task-relevant strengths to improve a VLA model's performance on challenging tasks. Extensive experiments in both simulation and real-world environments, across diverse sources of guidance, demonstrate that OMNIGUIDE enhances the performance of state-of-the-art generalist policies (e.g., $π_{0.5}$, GR00T N1.6) significantly across success and safety rates. Critically, our unified framework matches or surpasses the performance of prior methods designed to incorporate specific sources of guidance into VLA policies. Project Page: $\href{https://omniguide.github.io/}{this \; url}$
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。