arXiv:2501.04693cs.ROcs.AI2025-01ICRA被引 56

用语言连接多模态传感器,让机器人在视觉遮挡时仍能精准操作。

Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

  • 用语言作为跨模态桥梁,实现视觉、触觉、声音的联合建模。
  • 零样本下完成多模态提示与组合推理,真实世界成功率提升超20%。
  • 适配扩散模型和大视觉-语言-动作模型,通用性强。

与世界交互是多感官体验:实现有效通用交互需利用所有可用模态,包括视觉、触觉和听觉,以弥补部分观测缺失。例如,当视觉被遮挡时(如伸手进包),机器人应依赖触觉和听觉。然而,当前最先进的通用机器人策略通常仅基于大规模数据集,从视觉和本体感知中预测动作。本文提出FuSe,一种新方法,通过自然语言作为跨模态共同语义基础,使在缺乏大规模数据的异构传感器上微调视觉运动通用策略成为可能。我们结合多模态对比损失与感官-语言生成损失,编码高层语义。在机器人操作任务中,展示其可在零样本设置下完成需要联合推理多模态的任务,如多模态提示、组合跨模态提示及对交互物体的描述。相同方法适用于多种通用策略,包括基于扩散的通用策略与大型视觉-语言-动作(VLA)模型。大量真实世界实验表明,相比所有基线,成功率提升超过20%。

原文摘要 · Abstract (English)

Interacting with the world is a multi-sensory experience: achieving effective general-purpose interaction requires making use of all available modalities -- including vision, touch, and audio -- to fill in gaps from partial observation. For example, when vision is occluded reaching into a bag, a robot should rely on its senses of touch and sound. However, state-of-the-art generalist robot policies are typically trained on large datasets to predict robot actions solely from visual and proprioceptive observations. In this work, we propose FuSe, a novel approach that enables finetuning visuomotor generalist policies on heterogeneous sensor modalities for which large datasets are not readily available by leveraging natural language as a common cross-modal grounding. We combine a multimodal contrastive loss with a sensory-grounded language generation loss to encode high-level semantics. In the context of robot manipulation, we show that FuSe enables performing challenging tasks that require reasoning jointly over modalities such as vision, touch, and sound in a zero-shot setting, such as multimodal prompting, compositional cross-modal prompting, and descriptions of objects it interacts with. We show that the same recipe is applicable to widely different generalist policies, including both diffusion-based generalist policies and large vision-language-action (VLA) models. Extensive experiments in the real world show that FuSeis able to increase success rates by over 20% compared to all considered baselines.

机器人多模态语言接地零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。