用开放高斯点扩展视域,从少视角重建语义一致的3D场景。
OGGSplat: Open Gaussian Growing for Generalizable Reconstruction with Expanded Field-of-View
- 基于开放高斯点,通过图像扩散与语义扩散双向控制补全视野外区域。
- 在仅两视角输入下,手机拍摄图像也能实现语义感知的3D重建。
- 提出新基准GO,评估开放词汇场景的语义与生成质量。
从稀疏视图重建语义感知的3D场景是虚拟现实与具身AI等新兴应用的关键挑战。现有逐场景优化方法需密集输入视图且计算成本高,而通用方法常难以重建输入视锥外区域。本文提出OGGSplat,一种开放高斯生长方法,用于提升通用3D重建的视域范围。核心思想是:开放高斯的语义属性可为图像外推提供强先验,保障语义一致性与视觉合理性。具体地,从稀疏视图初始化开放高斯后,对选定渲染视图应用RGB-语义一致的图像修复模块,该模块在图像扩散模型与语义扩散模型间建立双向控制。修复区域被回传至3D空间,实现高效渐进的高斯参数优化。为评估方法,我们构建了高斯外推(Gaussian Outpainting, GO)基准,用于衡量开放词汇场景的语义与生成质量。实验表明,仅需两幅智能手机拍摄视图,OGGSplat即具备出色的语义感知场景重建能力。
原文摘要 · Abstract (English)
Reconstructing semantic-aware 3D scenes from sparse views is a challenging yet essential research direction, driven by the demands of emerging applications such as virtual reality and embodied AI. Existing per-scene optimization methods require dense input views and incur high computational costs, while generalizable approaches often struggle to reconstruct regions outside the input view cone. In this paper, we propose OGGSplat, an open Gaussian growing method that expands the field-of-view in generalizable 3D reconstruction. Our key insight is that the semantic attributes of open Gaussians provide strong priors for image extrapolation, enabling both semantic consistency and visual plausibility. Specifically, once open Gaussians are initialized from sparse views, we introduce an RGB-semantic consistent inpainting module applied to selected rendered views. This module enforces bidirectional control between an image diffusion model and a semantic diffusion model. The inpainted regions are then lifted back into 3D space for efficient and progressive Gaussian parameter optimization. To evaluate our method, we establish a Gaussian Outpainting (GO) benchmark that assesses both semantic and generative quality of reconstructed open-vocabulary scenes. OGGSplat also demonstrates promising semantic-aware scene reconstruction capabilities when provided with two view images captured directly from a smartphone camera.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。