将语言指令中的空间信息分离提取,提升机器人在少样本下的精准操作能力。
Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

- 用视觉语言模型生成零样本几何掩码,构建空间视觉提示。
- 50~100次演示下任务成功率从24.0%提升至39.5%,标准基准达67.8%。
- 适合少样本、高歧义语言指令的机器人操作场景。
端到端视觉-语言-动作(VLA)模型虽在机器人操作中展现潜力,但其整体架构将语义推理与空间控制耦合,造成严重对齐瓶颈,限制了数据受限下的精准目标识别。为此,我们提出SVP-IL,一种解耦架构,显式从动作生成流程中提取空间视觉定位。通过视觉语言基础模型,将指令解析为零样本几何掩码,实现语言到显式空间视觉提示(SVP)的转换。这些先验通过轻量级特征级融合机制注入连续动作生成器,提供明确且未污染的空间梯度引导,在低数据条件下仍保持高度稳定优化。大量实验表明,SVP-IL显著优于现有先进VLAs及纯视觉运动基线。在仅50至100次示范下,其在高度歧义语言条件任务上的平均成功率从24.0%提升至39.5%,标准基准达到67.8%。真实机器人实验进一步验证其在非结构化物理环境中的鲁棒性与数据效率。
原文摘要 · Abstract (English)
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This creates a severe alignment bottleneck, limiting precise target disambiguation in data-constrained imitation learning. To overcome this, we propose SVP-IL, a decoupled architecture that explicitly extracts spatial visual grounding from the action generation loop. By leveraging vision-language foundation models, we parse instructions into zero-shot geometric masks, translating language into explicit Spatial Visual Prompts (SVP). These priors are injected into a continuous action generator via a lightweight direct feature-level fusion mechanism. This integration provides explicit and uncorrupted spatial gradient guidance while ensuring highly stable optimization under low-data regimes. Extensive experiments demonstrate that SVP-IL significantly outperforms state-of-the-art VLAs and pure visuomotor baselines. Trained on as few as 50 to 100 demonstrations, SVP-IL improves average success rates on highly ambiguous language-conditioned tasks from 24.0% to 39.5%, achieving 67.8% on standard benchmarks. Real-world robotic experiments further validate its robustness and data efficiency in unstructured physical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。