用视觉语言模型实现机器人操作的多阶段引导,提升任务成功率。
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
- 通过空间与语义一致性微调,让模型理解任务细节
- 将任务拆解为多阶段子任务,提升轨迹敏感度
- 在稀疏奖励场景下表现更优,适合复杂操作任务
设计密集奖励函数对高效机器人强化学习至关重要。然而,大多数密集奖励依赖人工工程,从根本上限制了强化学习的可扩展性与自动化。尽管视觉语言模型(VLMs)为奖励设计提供了新路径,但原始的VLM奖励常与任务进展不一致,存在空间定位困难和任务语义理解不足的问题。为此,我们提出MARVL——基于视觉语言模型的机器人操作多阶段引导方法。MARVL通过对VLM进行空间与语义一致性微调,并将任务分解为多阶段子任务,结合任务方向投影以增强轨迹敏感性。实验表明,MARVL在Meta-World基准上显著优于现有VLM奖励方法,在稀疏奖励的操纵任务中展现出更优的样本效率与鲁棒性。
原文摘要 · Abstract (English)
Designing dense reward functions is pivotal for efficient robotic Reinforcement Learning (RL). However, most dense rewards rely on manual engineering, which fundamentally limits the scalability and automation of reinforcement learning. While Vision-Language Models (VLMs) offer a promising path to reward design, naive VLM rewards often misalign with task progress, struggle with spatial grounding, and show limited understanding of task semantics. To address these issues, we propose MARVL-Multi-stAge guidance for Robotic manipulation via Vision-Language models. MARVL fine-tunes a VLM for spatial and semantic consistency and decomposes tasks into multi-stage subtasks with task direction projection for trajectory sensitivity. Empirically, MARVL significantly outperforms existing VLM-reward methods on the Meta-World benchmark, demonstrating superior sample efficiency and robustness on sparse-reward manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。