arXiv:2502.12320cs.ROcs.CV2025-02被引 11

融合点云与图像信息,提升机器人模仿学习的感知能力

Towards Fusing Point Cloud and Visual Representations for Imitation Learning

  • 用全局与局部图像令牌条件化点云编码器,实现跨模态信息融合
  • 在RoboCasa基准上超越单一模态方法,全任务表现达最新水平
  • 适合需要精细几何与语义理解的机器人操作任务研究者

操作任务的学习需要能够访问丰富感官信息(如点云或RGB图像)的策略。点云能高效捕捉几何结构,对模仿学习中的操作任务至关重要;而RGB图像则提供丰富的纹理和语义信息,对某些任务尤为关键。现有融合方法将2D图像特征映射到点云,但常损失原始图像的全局上下文信息。本文提出FPV-Net,一种新型模仿学习方法,有效结合点云与RGB模态的优势。该方法通过自适应层归一化条件化,将全局和局部图像令牌融入点云编码器,充分利用两者的有益特性。在具有挑战性的RoboCasa基准上,实验表明仅依赖任一模态存在局限,而本文方法在所有任务中均达到当前最优性能。

原文摘要 · Abstract (English)

Learning for manipulation requires using policies that have access to rich sensory information such as point clouds or RGB images. Point clouds efficiently capture geometric structures, making them essential for manipulation tasks in imitation learning. In contrast, RGB images provide rich texture and semantic information that can be crucial for certain tasks. Existing approaches for fusing both modalities assign 2D image features to point clouds. However, such approaches often lose global contextual information from the original images. In this work, we propose FPV-Net, a novel imitation learning method that effectively combines the strengths of both point cloud and RGB modalities. Our method conditions the point-cloud encoder on global and local image tokens using adaptive layer norm conditioning, leveraging the beneficial properties of both modalities. Through extensive experiments on the challenging RoboCasa benchmark, we demonstrate the limitations of relying on either modality alone and show that our method achieves state-of-the-art performance across all tasks.

模仿学习多模态融合点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。