用视频插帧增强神经渲染,让3D重建更逼真
PS4PRO: Pixel-to-pixel Supervision for Photorealistic Rendering and Optimization
- 用视频插帧生成新视角,弥补真实拍摄角度缺失
- 在静态与动态场景中均提升重建质量,细节更清晰
- 适合做3D重建、虚拟拍摄的开发者和研究者
神经渲染方法通过多视图输入优化三维场景重建,但受限于视角数量。在复杂动态场景中,某些物体角度始终无法观测。本文提出以视频帧插值作为数据增强手段,并设计轻量级高质量插值模型PS4PRO(Pixel-to-pixel Supervision for Photorealistic Rendering and Optimization)。该模型在多样化视频数据集上训练,隐式建模相机运动与真实三维几何结构,作为隐式世界先验,增强对3D重建的图像监督。通过该方法,有效扩充了神经渲染所用数据集。实验表明,该方法在静态与动态场景中均显著提升重建性能。
原文摘要 · Abstract (English)
Neural rendering methods have gained significant attention for their ability to reconstruct 3D scenes from 2D images. The core idea is to take multiple views as input and optimize the reconstructed scene by minimizing the uncertainty in geometry and appearance across the views. However, the reconstruction quality is limited by the number of input views. This limitation is further pronounced in complex and dynamic scenes, where certain angles of objects are never seen. In this paper, we propose to use video frame interpolation as the data augmentation method for neural rendering. Furthermore, we design a lightweight yet high-quality video frame interpolation model, PS4PRO (Pixel-to-pixel Supervision for Photorealistic Rendering and Optimization). PS4PRO is trained on diverse video datasets, implicitly modeling camera movement as well as real-world 3D geometry. Our model performs as an implicit world prior, enriching the photo supervision for 3D reconstruction. By leveraging the proposed method, we effectively augment existing datasets for neural rendering methods. Our experimental results indicate that our method improves the reconstruction performance on both static and dynamic scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。