用3D高斯表示实现驾驶视频中物体精准可控的编辑
Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation
- 以3D高斯为稠密先验,提升生成时姿态控制精度
- 在Waymo数据集上实现物体重定位、插入与删除,视觉质量更优
- 适合自动驾驶场景合成与数据增强任务
罕见场景对自动驾驶系统的训练与验证至关重要,但真实采集成本高且危险。通过编辑已捕获传感器数据中的物体,可有效生成多样化场景,现有方法多依赖3D高斯点云或图像生成模型,但常存在视觉保真度不足或姿态控制不精确的问题。为此,我们提出G^2Editor框架,实现驾驶视频中物体的逼真且精确编辑。该方法利用待编辑物体的3D高斯表示作为稠密先验,注入去噪过程以确保姿态控制准确与空间一致性;采用场景级3D边界框布局重建非目标物体的遮挡区域;同时引入分层细粒度特征作为生成过程中的附加条件,指导编辑物体的外观细节。在Waymo Open Dataset上的实验表明,G^2Editor可在统一框架下实现物体重定位、插入与删除,显著优于现有方法,在姿态可控性与视觉质量方面表现更优,同时有益于下游数据驱动任务。
原文摘要 · Abstract (English)
Corner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for generating diverse scenarios, commonly achieved through 3D Gaussian Splatting or image generative models. However, these approaches often suffer from limited visual fidelity or imprecise pose control. To address these issues, we propose G^2Editor, a framework designed for photorealistic and precise object editing in driving videos. Our method leverages a 3D Gaussian representation of the edited object as a dense prior, injected into the denoising process to ensure accurate pose control and spatial consistency. A scene-level 3D bounding box layout is employed to reconstruct occluded areas of non-target objects. Furthermore, to guide the appearance details of the edited object, we incorporate hierarchical fine-grained features as additional conditions during generation. Experiments on the Waymo Open Dataset demonstrate that G^2Editor effectively supports object repositioning, insertion, and deletion within a unified framework, outperforming existing methods in both pose controllability and visual quality, while also benefiting downstream data-driven tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。