用真实车辆数据微调扩散模型,显著提升真实场景视角合成效果。
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
- 通过虚拟相机旋转对齐真实图像与合成数据几何结构。
- 在潜空间中进行遮挡感知训练,解决真实场景遮挡问题。
- 适合自动驾驶中车辆资产生成与3D视觉任务研究者使用。
大规模3D数据(如Objaverse)推动了姿态条件扩散模型在新视角合成上的进展,但其在真实图像上的表现明显下降。本文针对自动驾驶应用中的真实车辆资产采集,提出一系列微调策略。首先,通过虚拟相机旋转使真实图像与合成数据在几何上对齐,并保持与预训练模型定义的姿态流形一致;其次,采用以固定焦距学习不同物体尺度的物体中心化数据整理方式,应对真实驾驶场景中物体距离变化;此外,在潜空间中引入遮挡感知训练,以处理真实数据中普遍存在的遮挡问题,并利用对称先验处理大视角变化。这些方法有效提升了模型性能,使新视角合成的FID相比之前方法降低68.8%。
原文摘要 · Abstract (English)
The recent advent of large-scale 3D data, e.g. Objaverse, has led to impressive progress in training pose-conditioned diffusion models for novel view synthesis. However, due to the synthetic nature of such 3D data, their performance drops significantly when applied to real-world images. This paper consolidates a set of good practices to finetune large pretrained models for a real-world task -- harvesting vehicle assets for autonomous driving applications. To this end, we delve into the discrepancies between the synthetic data and real driving data, then develop several strategies to account for them properly. Specifically, we start with a virtual camera rotation of real images to ensure geometric alignment with synthetic data and consistency with the pose manifold defined by pretrained models. We also identify important design choices in object-centric data curation to account for varying object distances in real driving scenes -- learn across varying object scales with fixed camera focal length. Further, we perform occlusion-aware training in latent spaces to account for ubiquitous occlusions in real data, and handle large viewpoint changes by leveraging a symmetric prior. Our insights lead to effective finetuning that results in a $68.8\%$ reduction in FID for novel view synthesis over prior arts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。