通过模块化设计实现零样本长程操作,提升机器人任务鲁棒性。
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
- 分拆移动与交互模块,对象中心策略增强环境适应性
- 仿真中平均成功率69%,实测达85%,优于现有方法41%-67%
- 适合需要动态容错和技能复用的复杂现实场景
通用机器人需掌握长程操作,即在非结构化环境中涉及多个运动结构变化(如物体连接或分离)的任务。尽管视觉-语言-动作(VLA)模型具备掌握多样化基础技能的潜力,但在组合序列规划方面仍面临巨大挑战,且易受环境干扰导致级联失败。为此,我们提出LiLo-VLA(链接局部VLA),一种模块化框架,可在未训练过的新型长程任务上实现零样本泛化。该方法将全局运动与交互解耦:移动模块负责整体位移,交互模块采用对象中心VLA处理目标物体,确保对无关视觉特征的鲁棒性及空间配置不变性。关键在于,模块化设计支持动态重规划与技能复用,有效缓解端到端方法常见的级联错误。我们构建了包含21个任务的仿真基准,涵盖两个挑战性子集:LIBERO-Long++ 和 Ultra-Long。在仿真中,LiLo-VLA平均成功率达69%,比Pi0.5高41%,比OpenVLA-OFT高67%。真实世界评估在8个长程任务中平均成功率为85%。
原文摘要 · Abstract (English)
General-purpose robots must master long-horizon manipulation, defined as tasks involving multiple kinematic structure changes (e.g., attaching or detaching objects) in unstructured environments. While Vision-Language-Action (VLA) models offer the potential to master diverse atomic skills, they struggle with the combinatorial complexity of sequencing them and are prone to cascading failures due to environmental sensitivity. To address these challenges, we propose LiLo-VLA (Linked Local VLA), a modular framework capable of zero-shot generalization to novel long-horizon tasks without ever being trained on them. Our approach decouples transport from interaction: a Reaching Module handles global motion, while an Interaction Module employs an object-centric VLA to process isolated objects of interest, ensuring robustness against irrelevant visual features and invariance to spatial configurations. Crucially, this modularity facilitates robust failure recovery through dynamic replanning and skill reuse, effectively mitigating the cascading errors common in end-to-end approaches. We introduce a 21-task simulation benchmark consisting of two challenging suites: LIBERO-Long++ and Ultra-Long. In these simulations, LiLo-VLA achieves a 69% average success rate, outperforming Pi0.5 by 41% and OpenVLA-OFT by 67%. Furthermore, real-world evaluations across 8 long-horizon tasks demonstrate an average success rate of 85%. Project page: https://yy-gx.github.io/LiLo-VLA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。