arXiv:2604.02583cs.CV2026-04

融合多视角图像与3D模型,提升跨模态检索准确率

FusionBERT: Multi-View Image--3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder

  • 用交叉注意力融合多视角图像特征,增强视觉表示
  • 在真实工业物体上实现比现有模型更高的检索精度
  • 适合做3D模型检索、多视角视觉建模的研究者

我们提出FusionBERT,一种用于图像-3D跨模态检索的新型多视角视觉融合框架。现有方法主要关注单张物体图像与3D模型之间的特征对齐,难以应对实际场景中物体从多个视角观测的情况。尽管多视角能提供互补的几何与外观信息,但现有大模型很少有效融合此类信息。为此,FusionBERT引入基于交叉注意力的多视角视觉聚合器,自适应整合物体多视角图像特征,融合视间互补关系并强调关键视觉线索,生成更鲁棒的融合特征。此外,该模型还设计了法向感知的3D编码器,联合编码点法向量与3D位置,增强无纹理或颜色退化3D模型的几何表征。在合成3D模型和真实工业机械物体上的大量实验表明,FusionBERT在单视角与多视角设置下均显著优于当前最优多模态大模型,为多视角跨模态检索建立了新基准。

原文摘要 · Abstract (English)

We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model, limiting their applicability in realistic scenarios where an object is typically observed and captured from multiple viewpoints. Although multi-view observations naturally provide complementary geometric and appearance cues, existing multimodal large models rarely explore how to effectively fuse such multi-view visual information for better cross-modal retrieval. To address this limitation, we introduce a multi-view image--3D retrieval framework named FusionBERT, which innovatively utilizes a cross-attention-based multi-view visual aggregator to adaptively integrate features from multi-view images of an object. The proposed multi-view visual encoder fuses inter-view complementary relationships and selectively emphasizes informative visual cues across multiple views to get a more robustly fused visual feature for better 3D model matching. Furthermore, FusionBERT proposes a normal-aware 3D model encoder that can further enhance the 3D geometric feature of an object model by jointly encoding point normals and 3D positions, enabling a more robust representation learning for textureless or color-degraded 3D models. Extensive image--3D retrieval experiments on both synthetic 3D models and real-world industrial mechanical objects demonstrate that FusionBERT achieves significantly higher retrieval accuracy than SOTA multimodal large models under both single-view and multi-view settings, establishing a strong baseline for multi-view multimodal retrieval.

3D检索多视角融合视觉编码跨模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。