arXiv:2604.12551cs.CV2026-04

通过跨视图注意力融合多视角视觉语言嵌入,提升3D语义分割性能。

Cross-Attentive Multiview Fusion of Vision-Language Embeddings

论文配图:Cross-Attentive Multiview Fusion of Vision-Language Embeddings
图 1 · 摘自论文原文
  • 设计跨视图注意力机制,统一融合多视角视觉语言特征。
  • 引入多视图一致性自监督信号,显著提升分类准确率。
  • 适用于3D场景理解,尤其在零样本迁移上表现优异。

视觉语言模型推动了开放词汇2D语义分割的发展,但将其从2D图像扩展到3D场景仍具挑战。现有方法通常将2D描述符反投影并平均或启发式选择单一代表性描述符,常导致次优的3D表示。本文提出一种新型多视图变压器架构——跨注意力多视图融合(CAMFusion),通过跨视图注意力机制对多视角视觉语言嵌入进行融合,生成统一的每3D实例嵌入。其次,我们利用多视图一致性作为自监督信号,结合标准监督类别损失,显著提升性能。CAMFusion不仅持续优于简单平均或单视图选择,还在3D语义与实例分类基准上取得当前最优结果,包括在域外数据集上的零样本评估表现。

原文摘要 · Abstract (English)

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challenging problem. Existing approaches typically back-project and average 2D descriptors across views, or heuristically select a single representative one, often resulting in suboptimal 3D representations. In this work, we introduce a novel multiview transformer architecture that cross-attends across vision-language descriptors from multiple viewpoints and fuses them into a unified per-3D-instance embedding. As a second contribution, we leverage multiview consistency as a self-supervision signal for this fusion, which significantly improves performance when added to a standard supervised target-class loss. Our Cross-Attentive Multiview Fusion, which we denote with its acronym CAMFusion, not only consistently outperforms naive averaging or single-view descriptor selection, but also achieves state-of-the-art results on 3D semantic and instance classification benchmarks, including zero-shot evaluations on out-of-domain datasets.

3D分割多视图融合视觉语言模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。