让3D高斯点云在少视角下也能补全结构,靠语义信息精准修复缺失区域。
Seeing the Unseen: Semantic-in-Gaussian for Sparse-View 3D Generalization

- 用语义信息指导高斯点更新,实现跨视图一致性修复
- 在极限外推场景下PSNR提升2.44dB,结构更完整
- 适合少视角3D重建与复杂遮挡场景的科研人员
可泛化的3D高斯泼溅(G-3DGS)已成为稀疏视角下新视角合成的有前景方法。然而,现有框架受限于像素对齐的高斯估计,在部分观测或遮挡区域表现不佳,常导致表面不完整或结构坍塌。为此,我们提出SeeU(Seeing the Unseen)——一种新型G-3DGS框架。其核心设计为「语义嵌入高斯」:利用跨视图熵感知(CEA)模块,将多视角语义与几何线索融合为紧凑嵌入,引导条件高斯变压器对粗略高斯进行残差更新,有效恢复部分观测结构中的欠约束区域,同时保持表面一致性。在多个基准上的全面实验表明,SeeU在渲染质量与结构完整性上均持续提升,且维持高效前向推理。尤其在挑战性外推设置下,相比近期SOTA G-3DGS方法,平均PSNR提升2.44 dB。
原文摘要 · Abstract (English)
Generalizable 3D Gaussian Splatting (G-3DGS) has emerged as a promising approach for novel view synthesis undersparse-view settings. However, existing frameworks remain restricted by pixel-aligned Gaussian estimation, whichstruggles in partially observed or occluded regions and often leads to incomplete surfaces or structural collapse. Toaddress these challenges, we propose SeeU (Seeing the Unseen), a novel G-3DGS framework. We frame its core design asSemantic-in-Gaussian: semantic-conditioned refinement in Gaussian space. Specifically, we introduce a Cross-viewEntropy-Aware (CEA) module that aggregates multi-view semantic and geometric cues into compact embeddings. Theseembeddings guide the Conditional Gaussian Transformer, which applies residual updates to coarse Gaussians, helpingrecover under-constrained regions of partially observed structures while preserving surface consistency. Comprehensiveexperiments on multiple benchmarks demonstrate that SeeU consistently improves rendering quality and structuralcompleteness while retaining efficient feed-forward inference. Especially under challenging extrapolation settings,SeeU achieves an average improvement of 2.44 dB in PSNR compared to recent SOTA G-3DGS methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。