arXiv:2503.16413cs.CVcs.RO2025-03ICLR被引 3

M3让机器人记住复杂场景,用3D高斯点云存多模态信息。

M3: 3D-Spatial MultiModal Memory

  • 用3D高斯点云+注意力机制,高效存储多模态特征。
  • 在多个基础模型上实现90%以上特征相似度,支持跨模态理解。
  • 首次解决3D特征压缩中的对齐与失真问题,适合机器人视觉应用。

我们提出3D空间多模态记忆(M3),一种通过视频源保留中等规模静态场景信息的多模态记忆系统。结合3D高斯点阵技术和基础模型,M3构建了可跨粒度渲染特征表示的记忆体,涵盖广泛知识。研究发现先前特征点阵工作存在两大挑战:(1) 存储每个高斯基元的高维特征导致计算受限;(2) 提炼特征与基础模型特征间出现错位或信息丢失。为此,我们提出M3,包含主场景组件和高斯记忆注意力机制,实现高效训练与推理。通过全面的定量评估(特征相似度、下游任务)及定性可视化,验证其性能。该方法兼容多种基础模型,包括视觉-语言模型(VLM)、感知模型、大型多模态与语言模型(LMMs/LLMs)。此外,我们在四足机器人上部署了室内场景的M3特征场,证明其实用性。值得注意的是,M3是首个解决3D特征蒸馏核心压缩挑战的工作。

原文摘要 · Abstract (English)

We present 3D Spatial MultiModal Memory (M3), a multimodal memory system designed to retain information about medium-sized static scenes through video sources for visual perception. By integrating 3D Gaussian Splatting techniques with foundation models, M3 builds a multimodal memory capable of rendering feature representations across granularities, encompassing a wide range of knowledge. In our exploration, we identify two key challenges in previous works on feature splatting: (1) computational constraints in storing high-dimensional features for each Gaussian primitive, and (2) misalignment or information loss between distilled features and foundation model features. To address these challenges, we propose M3 with key components of principal scene components and Gaussian memory attention, enabling efficient training and inference. To validate M3, we conduct comprehensive quantitative evaluations of feature similarity and downstream tasks, as well as qualitative visualizations to highlight the pixel trace of Gaussian memory attention. Our approach encompasses a diverse range of foundation models, including vision-language models (VLMs), perception models, and large multimodal and language models (LMMs/LLMs). Furthermore, to demonstrate real-world applicability, we deploy M3's feature field in indoor scenes on a quadruped robot. Notably, we claim that M3 is the first work to address the core compression challenges in 3D feature distillation.

多模态记忆3D高斯机器人感知特征蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。