让多视角图像生成更稳定的3D世界表示,提升新视角合成效果
Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

- 用统一坐标系融合多视角深度、语义和置信度信息
- 新视角合成峰值信噪比最高提升2.5分贝,语义分割准确率提高11%
- 适合需要稳定3D表征的视觉任务,如机器人感知与重建
前馈高斯点云(FFGS)实现了实时新视角合成,但现有方法常依赖特定视角预测。随着输入视角增多,噪声或冗余信息会累积,难以收敛到稳定场景表征。本文提出CanonicalGS,一种前馈管道,将杂乱的多视角观测映射为稳定、以场景为中心的表示。该方法首先从深度、语义特征和不确定性估计中提取视角相关证据,再通过基于不确定性的融合机制,在一个规范的潜在世界中聚合这些证据。通过强化可靠观测、抑制不确定或冗余信息,CanonicalGS在新视角合成中表现更优,并可迁移至下游视觉任务。实验表明,新视角合成的峰值信噪比最高提升2.5 dB,语义分割准确率提升11%。
原文摘要 · Abstract (English)
Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are added, they may accumulate noisy or redundant evidence instead of converging to a stable scene representation. In this paper, we introduce CanonicalGS, a feed-forward pipeline that maps cluttered multi-view observations into a stable, scene-centric representation. CanonicalGS first extracts view-centric evidence from depth, semantic features, and uncertainty estimates, and then aggregates this evidence in a canonical latent world using uncertainty-aware fusion. By emphasizing reliable observations while suppressing uncertain or redundant ones, CanonicalGS produces representations that scale more effectively for novel view synthesis and transfer to downstream visual perception tasks. Experiments show up to a $2.5$ dB improvement in peak signal-to-noise ratio for synthesizing novel views and an $11\%$ gain in semantic segmentation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。