用三维重建增强数据集,提升大场景新视角生成效果
Aug3D: Augmenting large scale outdoor datasets for Generalizable Novel View Synthesis
- 基于SfM重建场景,通过网格与语义采样生成新视图
- 视图数从20减至10使PSNR提升10%,但性能仍不足
- 新生成视图与原数据融合,显著改善模型泛化能力
近期的逼真新视角生成(NVS)技术备受关注,但主要局限于小规模室内场景。尽管基于优化的方法已尝试解决此问题,具备显著优势的可泛化前馈方法仍研究不足。本文在大规模室外数据集UrbanScene3D上训练了PixelNeRF这一前馈NVS模型,提出四种训练策略,指出性能受限于视图重叠度不足。为此,我们引入Aug3D增强技术,利用传统运动恢复结构(SfM)重建场景,通过网格与语义采样生成高质量新视图,以提升前馈NVS模型学习效果。实验表明,每簇视图数从20降至10时,PSNR提升10%,但性能仍不理想;而将生成的新视图与原始数据结合后,显著增强了模型预测新视图的能力。
原文摘要 · Abstract (English)
Recent photorealistic Novel View Synthesis (NVS) advances have increasingly gained attention. However, these approaches remain constrained to small indoor scenes. While optimization-based NVS models have attempted to address this, generalizable feed-forward methods, offering significant advantages, remain underexplored. In this work, we train PixelNeRF, a feed-forward NVS model, on the large-scale UrbanScene3D dataset. We propose four training strategies to cluster and train on this dataset, highlighting that performance is hindered by limited view overlap. To address this, we introduce Aug3D, an augmentation technique that leverages reconstructed scenes using traditional Structure-from-Motion (SfM). Aug3D generates well-conditioned novel views through grid and semantic sampling to enhance feed-forward NVS model learning. Our experiments reveal that reducing the number of views per cluster from 20 to 10 improves PSNR by 10%, but the performance remains suboptimal. Aug3D further addresses this by combining the newly generated novel views with the original dataset, demonstrating its effectiveness in improving the model's ability to predict novel views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。