用机器人状态提升视觉语言与动作对齐,让机器人更懂自己、更少依赖人工数据。
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
- 通过自动获取的机器人状态数据,增强视觉语言模型对物理空间的理解。
- 在低数据场景下显著提升性能,减少对人类示范的依赖。
- 适合做少样本机器人控制、具身智能研究的团队参考。
视觉-语言-动作(VLA)模型在多任务端到端机器人控制中取得显著进展,得益于视觉-语言模型(VLM)的强大泛化能力。然而,其核心挑战在于如何有效对齐视觉语言空间与机器人动作空间。现有方法通常直接使用专家示范微调VLM,但存在时空鸿沟:空间上,VLM运行在高层语义空间,而机器人动作扎根于低层3D物理空间;时间上,VLM主要理解当前状态,而VLA模型需预测未来动作。为此,我们提出新型训练范式ROSA,利用机器人状态估计数据来改善视觉语言与动作空间的对齐。通过集成自动化获取的机器人状态信息,ROSA使VLA模型获得更强的空间理解力与自我认知能力,从而提升性能与泛化性。在模拟与真实环境中的大量实验表明,ROSA在低数据条件下尤为有效。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently made significant advance in multi-task, end-to-end robotic control, due to the strong generalization capabilities of Vision-Language Models (VLMs). A fundamental challenge in developing such models is effectively aligning the vision-language space with the robotic action space. Existing approaches typically rely on directly fine-tuning VLMs using expert demonstrations. However, this strategy suffers from a spatio-temporal gap, resulting in considerable data inefficiency and heavy reliance on human labor. Spatially, VLMs operate within a high-level semantic space, whereas robotic actions are grounded in low-level 3D physical space; temporally, VLMs primarily interpret the present, while VLA models anticipate future actions. To overcome these challenges, we propose a novel training paradigm, ROSA, which leverages robot state estimation to improve alignment between vision-language and action spaces. By integrating robot state estimation data obtained via an automated process, ROSA enables the VLA model to gain enhanced spatial understanding and self-awareness, thereby boosting performance and generalization. Extensive experiments in both simulated and real-world environments demonstrate the effectiveness of ROSA, particularly in low-data regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。