用3D高斯模型统一生成视角,提升无姿态图像的渲染质量。
SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis

- 复用同一3DGS场景的三种信号:几何、可见性、特征
- 在RealEstate10K上优于有真实姿态的基线方法
- 适合无姿态多视角图像生成任务的研究者
从无姿态图像生成逼真新视角,需兼具3D几何理解与未见内容合成能力。现有方法通常仅从3DGS重建中提取单一信号(像素渲染或学习特征),未利用每个高斯体的可见性信息进行遮挡感知的参考视图选择。这种‘信息断连’导致可渲染几何、可见性线索与学习特征均被浪费。SplatGuide通过将单个3DGS场景用于三个互补角色来弥合此断连:渲染图像提供像素对齐几何条件;每高斯体的源视图索引生成目标视图投票图,实现遮挡感知参考选择;重建令牌通过交叉注意力提供特征级引导。三类信号均来自一次前向传播。在RealEstate10K、DL3DV、Tanks-and-Temples和Mip-NeRF 360数据集上,SplatGuide达到无姿态新视角合成最优性能。在RealEstate10K上,使用中等数量输入视图时,超越基于真实姿态的基线。
原文摘要 · Abstract (English)
Generating photorealistic novel views from unposed images requires both 3D geometric understanding and the ability to synthesize unseen content. A natural strategy combines feed-forward 3DGS reconstruction with multi-view diffusion. Yet prior pipelines extract at most one signal from the reconstruction, either pixel rendering or learned features, while none exploits per-Gaussian visibility for occlusion-aware reference selection. This *information disconnect* leaves renderable geometry, visibility cues, and learned features unused. SplatGuide closes this disconnect by reusing a single 3DGS scene across three complementary roles. Rendered images provide pixel-aligned geometric conditioning. Per-Gaussian source-view indices are rendered into a target-view voting map for occlusion-aware reference selection. Reconstruction tokens supply feature-level guidance via cross-attention. All three signals derive from the same reconstruction forward pass. Across RealEstate10K, DL3DV, Tanks-and-Temples, and Mip-NeRF 360, SplatGuide achieves state-of-the-art pose-free novel view synthesis. On RealEstate10K, with a moderate number of input views, it surpasses the ground-truth-pose baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。