让机器人直接在视觉语言模型空间中决策,零样本完成新任务
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
- 反向对齐:将动作直接映射到VLM表示空间,保留细节信息
- 零样本生成最优闭环轨迹,实测在仿真和真实环境均超越基线
- 结合历史经验优化策略,适合需要快速适应新任务的机器人
视觉语言模型(VLM)在高维表示空间中编码了机器人操作的知识与推理能力。然而,现有方法常将其投影至压缩的中间表示,丢失细粒度的空间或语义信息。为此,我们提出AntiGrounding框架,逆转指令对齐过程:将候选动作直接提升至VLM表示空间,从多视角渲染轨迹,并通过结构化视觉问答实现基于指令的决策。该方法可实现新任务的零样本最优闭环轨迹合成。此外,我们设计离线策略精炼模块,利用过往经验提升长期性能。在仿真与真实环境中的实验表明,该方法在多种机器人操作任务中均优于基线。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) encode knowledge and reasoning capabilities for robotic manipulation within high-dimensional representation spaces. However, current approaches often project them into compressed intermediate representations, discarding important task-specific information such as fine-grained spatial or semantic details. To address this, we propose AntiGrounding, a new framework that reverses the instruction grounding process. It lifts candidate actions directly into the VLM representation space, renders trajectories from multiple views, and uses structured visual question answering for instruction-based decision making. This enables zero-shot synthesis of optimal closed-loop robot trajectories for new tasks. We also propose an offline policy refinement module that leverages past experience to enhance long-term performance. Experiments in both simulation and real-world environments show that our method outperforms baselines across diverse robotic manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。