用大模型零样本找3D场景相似处,让机器人跨环境复用经验。
Finding 3D Scene Analogies with Multimodal Foundation Models
- 用视觉语言模型和3D形状模型构建混合场景表示
- 在复杂场景间建立精确对应关系,支持轨迹与目标点迁移
- 无需训练、开放词汇,适合新环境快速适应的机器人应用
将当前观察与过往经验关联,有助于机器人在未见的3D环境中适应与规划。近期提出的3D场景类比通过平滑映射对齐具有共同空间关系的场景区域,实现轨迹或航点的精细传递,可支持模仿学习中的示范迁移或任务计划跨场景转移。然而,现有方法需额外训练且依赖固定物体词表。本文提出利用多模态基础模型,在零样本、开放词汇设置下寻找3D场景类比。核心是结合视觉语言模型特征的稀疏图与3D形状基础模型生成的特征场构成混合神经表示。类比关系以粗到精的方式建立:先对齐图结构,再用特征场细化对应关系。实验表明该方法能准确匹配复杂场景,并应用于轨迹与航点传递。
原文摘要 · Abstract (English)
Connecting current observations with prior experiences helps robots adapt and plan in new, unseen 3D environments. Recently, 3D scene analogies have been proposed to connect two 3D scenes, which are smooth maps that align scene regions with common spatial relationships. These maps enable detailed transfer of trajectories or waypoints, potentially supporting demonstration transfer for imitation learning or task plan transfer across scenes. However, existing methods for the task require additional training and fixed object vocabularies. In this work, we propose to use multimodal foundation models for finding 3D scene analogies in a zero-shot, open-vocabulary setting. Central to our approach is a hybrid neural representation of scenes that consists of a sparse graph based on vision-language model features and a feature field derived from 3D shape foundation models. 3D scene analogies are then found in a coarse-to-fine manner, by first aligning the graph and refining the correspondence with feature fields. Our method can establish accurate correspondences between complex scenes, and we showcase applications in trajectory and waypoint transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。