端到端统一重建与生成,让新视角视频更连贯
RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation

- 用隐式场景表示直接查询相机射线获取几何特征
- 在DL3DV上图像和视频一致性指标均超越基线
- 无需中间3D表示,联合训练提升生成质量
从稀疏输入进行新视角合成需兼顾观测视图的几何约束和未观测区域的生成先验,推动了融合重建与生成的混合方法发展。但现有方法通过渲染图像或显式3D表示(如点云、3D高斯)连接两者,导致生成依赖于有损且不完美的场景投影,继承其误差,而重建也缺乏生成信号来修正错误。本文提出RoGe,一个端到端统一的重建与生成框架,消除这种显式桥梁。该框架针对由稀疏视图锚定的场景内漫游:给定少数带姿态的图像和相机轨迹,合成沿该轨迹的时序一致视频。从稀疏输入视图出发,RoGe使用前馈重建模型构建隐式场景表示,并以目标相机射线查询得到每视图的几何特征。这些特征作为条件注入视频扩散模型,无需任何3D中间表示。两个模块联合训练,使生成目标直接塑造自身的几何条件。在DL3DV数据集上的实验表明,RoGe在图像级指标和视频级时间一致性上均优于基于重建、基于生成及混合基线。消融实验确认,射线查询的隐式特征优于原始重建令牌和渲染RGB作为条件,联合训练带来进一步提升。
原文摘要 · Abstract (English)
Novel view synthesis from sparse inputs requires both geometric grounding from the observed views and generative priors of unobserved regions, motivating recent hybrid methods that combine reconstruction and generation. However, existing methods bridge the two with rendered images or explicit 3D representations such as point maps or 3D Gaussians. Generation is thus conditioned on a lossy and imperfect projection of the scene, inheriting its errors, and reconstruction receives no signal from generation to correct them. We present RoGe, an end-to-end unified reconstruction and generation framework that removes this explicit bridge. It targets roaming within a scene anchored by sparse views: given a few posed images and a camera trajectory, it synthesizes a temporally coherent video along that trajectory. From the sparse input views, RoGe builds an implicit scene representation with a feed-forward reconstruction model, and queries it with target camera rays to obtain per-view geometric features. These features are injected into a video diffusion model as conditioning, without any 3D intermediate. Both modules are trained jointly, so the generation objective directly shapes its own geometric conditioning. We conduct experiments on DL3DV, where RoGe outperforms reconstruction-based, generation-based, and hybrid baselines on image-level metrics and video-level temporal consistency. Ablations confirm that ray-queried implicit features outperform both raw reconstruction tokens and rendered RGB as conditioning, and that joint training brings further gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。