arXiv:2501.16384cs.CVcs.LG2025-01中稿 · the Workshop on Im…被引 5

用Mamba提升点云补全效率,图像引导更省算力。

MambaTron: Efficient Cross-Modal Point Cloud Enhancement using Aggregate Selective State Space Modeling

  • 设计Mamba-Transformer混合单元,融合长序列高效与注意力优势
  • 在视图引导点云补全任务上达到顶尖性能,计算量仅为现有方法几分之一
  • 适合需要低延迟、高效率多模态3D重建的场景

点云增强旨在从不完整输入生成高质量点云,通过回归填补缺失细节。本文聚焦视图引导的点云补全任务,即利用图像(代表点云视角)中的信息来生成完整点云。尽管状态空间模型(如Mamba)在自然语言处理及2D/3D视觉中展现出对自注意力机制的高效替代潜力,但其在图像与点云间的跨模态注意力应用仍较少。为此,本文提出MambaTron——一种基于Mamba的Transformer单元,作为网络的基本构建模块,支持单模态与跨模态重建,包括视图引导点云补全。该方法结合Mamba的长序列处理效率与Transformer的强分析能力,是首个在计算机视觉中实现基于Mamba的跨注意力机制尝试。实验表明,模型性能接近当前最优水平,同时仅需极少量计算资源。

原文摘要 · Abstract (English)

Point cloud enhancement is the process of generating a high-quality point cloud from an incomplete input. This is done by filling in the missing details from a reference like the ground truth via regression, for example. In addition to unimodal image and point cloud reconstruction, we focus on the task of view-guided point cloud completion, where we gather the missing information from an image, which represents a view of the point cloud and use it to generate the output point cloud. With the recent research efforts surrounding state-space models, originally in natural language processing and now in 2D and 3D vision, Mamba has shown promising results as an efficient alternative to the self-attention mechanism. However, there is limited research towards employing Mamba for cross-attention between the image and the input point cloud, which is crucial in multi-modal problems. In this paper, we introduce MambaTron, a Mamba-Transformer cell that serves as a building block for our network which is capable of unimodal and cross-modal reconstruction which includes view-guided point cloud completion.We explore the benefits of Mamba's long-sequence efficiency coupled with the Transformer's excellent analytical capabilities through MambaTron. This approach is one of the first attempts to implement a Mamba-based analogue of cross-attention, especially in computer vision. Our model demonstrates a degree of performance comparable to the current state-of-the-art techniques while using a fraction of the computation resources.

点云补全多模态Mamba高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。