解决单目拍摄场景尺度模糊问题,提升生成新视角图像的一致性。
Learn Your Scales: Towards Scale-Consistent Generative Novel View Synthesis
- 联合训练生成模型与场景尺度估计,端到端消除尺度歧义。
- 新指标显示生成视角尺度不一致问题显著降低。
- 无需复杂预处理,提升生成图像质量,适合3D生成研究者。
传统无深度的多视角数据由未校准的单目相机移动拍摄获得,其相机位置尺度存在歧义。以往方法通过各种手动归一化处理应对尺度模糊,但未深入分析错误尺度对生成式新视角合成(GNVS)的影响。本文研究单图输入下尺度歧义对GNVS模型的影响,提出新度量标准以评估生成视角的尺度一致性,并设计端到端框架联合学习场景尺度与生成模型。实验表明,该方法在不引入复杂预处理的情况下,有效减少生成视角的尺度不一致,同时提升生成图像质量。
原文摘要 · Abstract (English)
Conventional depth-free multi-view datasets are captured using a moving monocular camera without metric calibration. The scales of camera positions in this monocular setting are ambiguous. Previous methods have acknowledged scale ambiguity in multi-view data via various ad-hoc normalization pre-processing steps, but have not directly analyzed the effect of incorrect scene scales on their application. In this paper, we seek to understand and address the effect of scale ambiguity when used to train generative novel view synthesis methods (GNVS). In GNVS, new views of a scene or object can be minimally synthesized given a single image and are, thus, unconstrained, necessitating the use of generative methods. The generative nature of these models captures all aspects of uncertainty, including any uncertainty of scene scales, which act as nuisance variables for the task. We study the effect of scene scale ambiguity in GNVS when sampled from a single image by isolating its effect on the resulting models and, based on these intuitions, define new metrics that measure the scale inconsistency of generated views. We then propose a framework to estimate scene scales jointly with the GNVS model in an end-to-end fashion. Empirically, we show that our method reduces the scale inconsistency of generated views without the complexity or downsides of previous scale normalization methods. Further, we show that removing this ambiguity improves generated image quality of the resulting GNVS model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。