arXiv:2604.02546cs.CVcs.LG2026-04中稿 · ed

用多视角图像点云预训练3D场景理解,提升泛化能力

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

论文配图:RGB-Pointmap Pretraining for Unified 3D Scene Understanding
图 1 · 摘自论文原文
  • 基于视觉-语言预训练模型,融合多视角图像与点云学习统一表征
  • 在少样本任务上达到当前最优,支持场景分类、问答等多任务
  • 适合做3D视觉理解的通用模型研究者或工业应用开发者

通过与对比语言-图像预训练(CLIP)对齐来预训练3D编码器,已成为学习可泛化3D场景理解表示的有前景方向。本文提出UniScene3D,一种基于Transformer的框架,通过利用预训练2D基础模型的先验知识,从多视角RGB-Pointmap输入中学习统一的3D场景表示。为实现鲁棒的RGB-Pointmap表示学习,引入跨视图几何对齐和基于语义的视图对齐,以保证视图间的几何与语义一致性。在视角定位、场景检索、场景分类和3D视觉问答等任务上的低样本及特定任务微调实验表明,该方法性能达到当前最优水平。这些结果证实了UniScene3D作为统一3D场景理解有效框架的潜力。

原文摘要 · Abstract (English)

Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding. In this paper, we propose UniScene3D, a transformer-based framework that learns unified 3D scene representations from multi-view RGB-Pointmap inputs by leveraging the priors of a pretrained 2D foundation model. For robust RGB-Pointmap representation learning, we introduce cross-view geometric alignment and grounded view alignment to enforce geometric and semantic consistency across views. Extensive low-shot and task-specific fine-tuning on viewpoint grounding, scene retrieval, scene classification, and 3D visual question answering demonstrates state-of-the-art performance. These results establish UniScene3D as an effective framework for unified 3D scene understanding. Project page: https://yebulabula.github.io/UniScene3D/

3D理解多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。