arXiv:2412.12906cs.CV2024-12ICCV被引 6

用文本和空间信息增强单视角3D高斯点云重建效果

CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image

  • 引入视觉语言模型文本引导与3D点特征空间指导
  • 单视角下实现高质量新视图合成,优于现有方法
  • 适合需要高效3D重建的场景,如AR/VR应用

近期基于3D高斯点云的前馈式通用方法因其在有限资源下重建3D场景的潜力而受到关注。这些方法仅需少量图像,在一次前向传播中即可构建由每像素3D高斯原语参数化的3D辐射场。然而,与依赖多视角对应关系的方法相比,单视角下的3D场景重建仍属研究空白。本文提出CATSplat,一种新型通用的基于Transformer的框架,旨在突破单目设定下的固有局限。首先,我们利用视觉-语言模型的文本提示来补充单张图像信息不足的问题。通过交叉注意力机制融合文本嵌入中的场景特定上下文细节,实现超越纯视觉线索的上下文感知3D重建。此外,我们倡导使用3D点特征提供的空间引导,以在单视角条件下获得更全面的几何理解。借助3D先验,图像特征可捕捉丰富的结构信息,从而在无需多视角技术的情况下预测3D高斯分布。大规模数据集上的大量实验表明,CATSplat在单视角3D场景重建中达到当前最佳性能,实现了高质量的新视图合成。

原文摘要 · Abstract (English)

Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. However, unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from a single image. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under single-view settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis.

3D重建单视角高斯点云Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。