用扩散模型精准编辑自动驾驶场景中的物体位置和外观。
DriveEditor: A Unified 3D Information-Guided Framework for Controllable Object Editing in Driving Scenes
- 基于3D边界框控制物体位置,融合深度信息实现精确定位。
- 通过三重策略保持物体外观一致,支持替换、删除等操作。
- 适用于自动驾驶数据增强,提升下游任务性能。
以视觉为中心的自动驾驶系统需要多样化的数据进行稳健训练与评估,可通过操纵现有场景捕获中的物体位置和外观来实现数据增强。尽管扩散模型在视频编辑方面取得进展,但在驾驶场景中进行物体操作仍面临位置控制不精确、高保真外观难以维持的挑战。为此,我们提出DriveEditor,一种基于扩散模型的驾驶视频物体编辑框架。DriveEditor提供统一框架,支持重新定位、替换、删除和插入等多种编辑操作,所有操作通过共享输入和相同的位置控制与外观保持模块完成。位置控制模块将给定3D边界框投影并保留深度信息,分层注入扩散过程,实现对物体位置和朝向的精确控制。外观保持模块通过三级方法:低层级细节保持、高层语义维持,以及利用新型视图合成模型的3D先验,确保单一参考图像下的属性一致性。在nuScenes数据集上的大量定性和定量评估表明,DriveEditor在生成多样化驾驶场景编辑时表现出卓越的保真度和可控性,并显著促进下游任务。项目页面:https://yvanliang.github.io/DriveEditor。
原文摘要 · Abstract (English)
Vision-centric autonomous driving systems require diverse data for robust training and evaluation, which can be augmented by manipulating object positions and appearances within existing scene captures. While recent advancements in diffusion models have shown promise in video editing, their application to object manipulation in driving scenarios remains challenging due to imprecise positional control and difficulties in preserving high-fidelity object appearances. To address these challenges in position and appearance control, we introduce DriveEditor, a diffusion-based framework for object editing in driving videos. DriveEditor offers a unified framework for comprehensive object editing operations, including repositioning, replacement, deletion, and insertion. These diverse manipulations are all achieved through a shared set of varying inputs, processed by identical position control and appearance maintenance modules. The position control module projects the given 3D bounding box while preserving depth information and hierarchically injects it into the diffusion process, enabling precise control over object position and orientation. The appearance maintenance module preserves consistent attributes with a single reference image by employing a three-tiered approach: low-level detail preservation, high-level semantic maintenance, and the integration of 3D priors from a novel view synthesis model. Extensive qualitative and quantitative evaluations on the nuScenes dataset demonstrate DriveEditor's exceptional fidelity and controllability in generating diverse driving scene edits, as well as its remarkable ability to facilitate downstream tasks. Project page: https://yvanliang.github.io/DriveEditor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。