通过解耦场景几何与外观,生成可编辑的高保真驾驶视角视频。
GA-Drive: Geometry-Appearance Decoupled Modeling for Free-viewpoint Driving Scene Generation
- 分离场景几何与外观,分步生成新视角图像。
- 在NTA-IoU、NTL-IoU和FID指标上显著优于现有方法。
- 支持外观编辑且保持几何一致性,适合自动驾驶训练模拟。
自由视角、可编辑且高保真的驾驶模拟器对端到端自动驾驶系统训练与评估至关重要。本文提出GA-Drive,一种新型模拟框架,通过几何-外观解耦与基于扩散的生成技术,沿用户指定的新轨迹生成相机视角。给定记录轨迹上的图像及对应场景几何信息,GA-Drive利用几何信息合成伪视角图像,再通过训练好的视频扩散模型将其转换为逼真视图。该方法实现几何与外观的解耦,优势在于可借助前沿视频到视频编辑技术进行外观修改,同时保持底层几何一致,确保原轨迹与新轨迹间编辑的一致性。大量实验表明,GA-Drive在NTA-IoU、NTL-IoU和FID评分上均显著优于现有方法。
原文摘要 · Abstract (English)
A free-viewpoint, editable, and high-fidelity driving simulator is crucial for training and evaluating end-to-end autonomous driving systems. In this paper, we present GA-Drive, a novel simulation framework capable of generating camera views along user-specified novel trajectories through Geometry-Appearance Decoupling and Diffusion-Based Generation. Given a set of images captured along a recorded trajectory and the corresponding scene geometry, GA-Drive synthesizes novel pseudo-views using geometry information. These pseudo-views are then transformed into photorealistic views using a trained video diffusion model. In this way, we decouple the geometry and appearance of scenes. An advantage of such decoupling is its support for appearance editing via state-of-the-art video-to-video editing techniques, while preserving the underlying geometry, enabling consistent edits across both original and novel trajectories. Extensive experiments demonstrate that GA-Drive substantially outperforms existing methods in terms of NTA-IoU, NTL-IoU, and FID scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。