arXiv:2504.01901cs.CVcs.AI2025-04ICCV被引 72

通过3D视觉监督提升模型对场景的重建能力,实现更精准的3D理解。

Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

  • 在训练中加入跨视角与全局视角重建任务,增强3D感知。
  • 在多个3D场景理解基准上达到顶尖性能,优于现有方法。
  • 可有效利用大量无标签3D视觉数据,适合大规模3D建模应用。

大型多模态模型(LMMs)在2D图像和视频领域的快速发展,推动了其向3D场景理解的迁移。然而,缺乏大规模3D视觉-语言数据集成为主要障碍。传统方法通过设计3D输入级场景表示来注入3D感知。本文提出一种新思路:将3D-aware视觉监督融入训练过程,构建名为Ross3D的可重构视觉指令微调框架。该方法包含两种重建机制:跨视角重建要求从其他视角聚合重叠信息以恢复被遮挡视图;全局视角重建则利用所有可用视角信息恢复鸟瞰图(Bird's-Eye-View),提供完整场景概览。实验表明,Ross3D在多个3D场景理解基准上达到当前最优性能。更重要的是,半监督实验显示其能有效利用大量未标注的3D视觉数据,展现出显著潜力。

原文摘要 · Abstract (English)

The rapid development of Large Multimodal Models (LMMs) for 2D images and videos has spurred efforts to adapt these models for interpreting 3D scenes. However, the absence of large-scale 3D vision-language datasets has posed a significant obstacle. To address this issue, typical approaches focus on injecting 3D awareness into 2D LMMs by designing 3D input-level scene representations. This work provides a new perspective. We introduce reconstructive visual instruction tuning with 3D-awareness (Ross3D), which integrates 3D-aware visual supervision into the training procedure. Specifically, it incorporates cross-view and global-view reconstruction. The former requires reconstructing masked views by aggregating overlapping information from other views. The latter aims to aggregate information from all available views to recover Bird's-Eye-View images, contributing to a comprehensive overview of the entire scene. Empirically, Ross3D achieves state-of-the-art performance across various 3D scene understanding benchmarks. More importantly, our semi-supervised experiments demonstrate significant potential in leveraging large amounts of unlabeled 3D vision-only data.

3D理解视觉重建多模态自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。