arXiv:2506.20991cs.CV2025-06AAAI

解决3D点云文本对齐难题,提升查询驱动分割精度

MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation

  • 设计显式跨模态对齐模块,直接连接3D点云与文本/图像
  • 引入视觉-文本记忆模块,动态召回场景相关语义特征
  • 在多个基准上达到领先性能,适合需要精准语义分割的场景

视觉语言模型在3D领域的快速发展推动了文本查询引导的点云处理研究,但现有方法在点级分割上表现不佳,主要受限于3D-文本对齐不足,难以建立局部特征与文本语义的关联。为此,我们提出MR-COSMO——一种用于查询驱动3D分割的视觉-文本记忆召回与直接跨模态对齐方法。通过专用的直接跨模态对齐模块,建立3D点云与文本/2D图像数据之间的显式对齐;同时引入视觉-文本记忆模块,配备专门的特征库,存储文本特征、视觉特征及其对应映射关系。该记忆模块通过基于注意力的知识召回机制,动态增强特定场景的表示能力。在3D指令、参考及语义分割等多个基准上的全面实验表明,该方法达到当前最优性能。

原文摘要 · Abstract (English)

The rapid advancement of vision-language models (VLMs) in 3D domains has accelerated research in text-query-guided point cloud processing, though existing methods underperform in point-level segmentation due to inadequate 3D-text alignment that limits local feature-text context linking. To address this limitation, we propose MR-COSMO, a Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation, establishing explicit alignment between 3D point clouds and text/2D image data through a dedicated direct cross-modal alignment module while implementing a visual-text memory module with specialized feature banks. This direct alignment mechanism enables precise fusion of geometric and semantic features, while the memory module employs specialized banks storing text features, visual features, and their correspondence mappings to dynamically enhance scene-specific representations via attention-based knowledge recall. Comprehensive experiments across 3D instruction, reference, and semantic segmentation benchmarks confirm state-of-the-art performance.

3D分割跨模态对齐视觉语言模型点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。