区分重建与生成区域,提升新视角合成的精度与真实感
GenRec: Knowing Where to Reconstruct and Where to Generate

- 用观察掩码划分可见与遮挡区域,分别进行重建与生成
- 在真实数据集上,可见区域重建误差更低,不可见区域视觉质量更优
- 适合追求高几何精度与自然幻觉平衡的研究者
从稀疏输入图像生成新视角图像时,部分像素(如可见区域)有唯一正确值,仅受视点相关光照影响;而遮挡或超出拍摄范围的像素则存在多种合理补全可能。现有方法使用统一损失混淆了重建与生成,即使引入几何信息也难以区分。本文提出GenRec,一种基于多视角流匹配的模型,将重建-生成分离直接嵌入架构、监督和梯度流中。通过源相机信息与单目深度估计获得观察掩码,流匹配主干联合去噪所有目标视图的RGB与场景坐标图,像素空间细化阶段恢复已知像素的高频细节;同一掩码控制监督,防止回归信号污染生成先验。在RealEstate10K、DL3DV-10K和Mip-NeRF~360数据集上,无论单视图外推还是双视图插值,GenRec在可见区域均实现最优重建保真度,同时在不可见区域的感知质量超越纯生成基线,验证了该方法的有效性。
原文摘要 · Abstract (English)
Generative novel view synthesis from sparse input images is rarely all reconstruction or all generation: pixels visible in some source view have a unique correct value modulated only by view-dependent shading, while pixels in disocclusions or beyond the captured volume admit a distribution of plausible completions. Existing generative novel-view-synthesis methods conflate these regimes under a single uniform loss, blurring the line between geometric fidelity and creative hallucinations even when scene geometry is injected through warped point clouds or projected depth. We introduce GenRec, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow. Guided by an observation mask derived from the source cameras and a monocular depth estimator, a flow matching backbone jointly denoises RGB and scene-coordinate maps across all target views, while a pixel-space refinement stage restores high-frequency detail on observed pixels; the same mask gates supervision so regression signals do not contaminate the generative prior. Across RealEstate10K, DL3DV-10K, and Mip-NeRF~360, in both single-view extrapolation and two-view interpolation, GenRec attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones, showing the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。