arXiv:2506.01600cs.ROcs.AI2025-06被引 8

无需专家示范,让机器人用语言指令在新环境中精准找物

WoMAP: World Models For Embodied Open-Vocabulary Object Localization

  • 用高斯点云构建真实-仿真-真实数据流,自动生成训练数据
  • 零样本定位成功率比视觉语言模型高9倍,比扩散策略高2倍
  • 适合需要快速适应新环境的机器人任务,尤其擅长跨场景迁移

语言指令驱动的主动物体定位是机器人面临的关键挑战,需高效探索部分可观测环境。现有方法或难以泛化到演示数据之外(如模仿学习),或无法生成物理上合理的动作(如视觉语言模型)。为此,我们提出WoMAP(用于主动感知的世界模型):一种训练开放词汇物体定位策略的方法,包含三部分:(i) 基于高斯点云的真实-仿真-真实数据生成管道,无需专家示范即可规模化生成数据;(ii) 从开放词汇物体检测器中蒸馏密集奖励信号;(iii) 利用潜在世界模型预测动态与奖励,在推理时将高层动作提议具身化。严格的仿真与硬件实验表明,WoMAP在多种零样本物体定位任务中表现卓越,其成功率分别比VLM和扩散策略基线高出9倍和2倍。此外,我们在TidyBot上验证了其强泛化能力与仿真到现实的迁移性能。

原文摘要 · Abstract (English)

Language-instructed active object localization is a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art approaches either struggle to generalize beyond demonstration datasets (e.g., imitation learning methods) or fail to generate physically grounded actions (e.g., VLMs). To address these limitations, we introduce WoMAP (World Models for Active Perception): a recipe for training open-vocabulary object localization policies that: (i) uses a Gaussian Splatting-based real-to-sim-to-real pipeline for scalable data generation without the need for expert demonstrations, (ii) distills dense rewards signals from open-vocabulary object detectors, and (iii) leverages a latent world model for dynamics and rewards prediction to ground high-level action proposals at inference time. Rigorous simulation and hardware experiments demonstrate WoMAP's superior performance in a broad range of zero-shot object localization tasks, with more than 9x and 2x higher success rates compared to VLM and diffusion policy baselines, respectively. Further, we show that WoMAP achieves strong generalization and sim-to-real transfer on a TidyBot.

机器人视觉语言模型零样本定位世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。