用合成图像提升极稀疏视角下全景图的新视角生成质量
Improving Novel view synthesis of 360$^\circ$ Scenes in Extremely Sparse Views by Jointly Training Hemisphere Sampled Synthetic Images
- 通过上半球密集采样生成合成图像,补充稀疏输入
- 结合真实与合成图像训练3D高斯泼溅模型,扩大场景覆盖
- 适配扩散模型增强渲染质量,适合虚拟现实应用
从极稀疏视角生成360°场景的新视角对虚拟现实和增强现实至关重要。传统结构光恢复方法在极稀疏情况下难以估计相机位姿,本文采用DUSt3R估计位姿并生成稠密点云。基于估计位姿,在场景上半球空间密集采样额外视角,结合点云渲染合成图像。将稀疏视角的真实图像与密集合成图像联合训练3D高斯泼溅模型,提升3D空间覆盖范围,缓解稀疏输入下的过拟合问题。进一步在自建数据集上微调基于扩散的图像增强模型,有效去除点云渲染中的伪影。在仅四张输入图像的极端稀疏条件下,相比基准方法,本框架显著提升了360°场景的新视角生成效果。
原文摘要 · Abstract (English)
Novel view synthesis in 360$^\circ$ scenes from extremely sparse input views is essential for applications like virtual reality and augmented reality. This paper presents a novel framework for novel view synthesis in extremely sparse-view cases. As typical structure-from-motion methods are unable to estimate camera poses in extremely sparse-view cases, we apply DUSt3R to estimate camera poses and generate a dense point cloud. Using the poses of estimated cameras, we densely sample additional views from the upper hemisphere space of the scenes, from which we render synthetic images together with the point cloud. Training 3D Gaussian Splatting model on a combination of reference images from sparse views and densely sampled synthetic images allows a larger scene coverage in 3D space, addressing the overfitting challenge due to the limited input in sparse-view cases. Retraining a diffusion-based image enhancement model on our created dataset, we further improve the quality of the point-cloud-rendered images by removing artifacts. We compare our framework with benchmark methods in cases of only four input views, demonstrating significant improvement in novel view synthesis under extremely sparse-view conditions for 360$^\circ$ scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。