arXiv:2507.21872cs.AI2025-07

用3D高斯点云先验实现驾驶场景下多模态物体可控编辑

MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors

  • 引入3D高斯点云作为目标物体的结构与外观先验
  • 多层级控制机制实现图像与激光点云的高保真重建
  • 适合需要生成罕见车辆数据以提升感知模型性能的研究者

自动驾驶系统依赖多模态感知数据理解复杂环境。然而,真实世界数据的长尾分布限制了模型泛化能力,尤其对罕见但关键的车辆类别。为此,我们提出MultiEditor,一种双分支潜在扩散框架,用于联合编辑驾驶场景中的图像与激光点云。核心是引入3D高斯点积(3DGS)作为目标物体的结构与外观先验。基于此先验,设计多层级外观控制机制——包括像素级粘贴、语义级引导和多分支优化——实现跨模态高保真重建。进一步提出深度引导的可变形跨模态条件模块,利用3DGS渲染深度自适应实现模态间相互指导,显著提升跨模态一致性。大量实验表明,MultiEditor在视觉与几何保真度、编辑可控性及跨模态一致性方面均表现优异。此外,使用MultiEditor生成罕见类别车辆数据,可显著提升感知模型对低频类别的检测准确率。

原文摘要 · Abstract (English)

Autonomous driving systems rely heavily on multimodal perception data to understand complex environments. However, the long-tailed distribution of real-world data hinders generalization, especially for rare but safety-critical vehicle categories. To address this challenge, we propose MultiEditor, a dual-branch latent diffusion framework designed to edit images and LiDAR point clouds in driving scenarios jointly. At the core of our approach is introducing 3D Gaussian Splatting (3DGS) as a structural and appearance prior for target objects. Leveraging this prior, we design a multi-level appearance control mechanism--comprising pixel-level pasting, semantic-level guidance, and multi-branch refinement--to achieve high-fidelity reconstruction across modalities. We further propose a depth-guided deformable cross-modality condition module that adaptively enables mutual guidance between modalities using 3DGS-rendered depth, significantly enhancing cross-modality consistency. Extensive experiments demonstrate that MultiEditor achieves superior performance in visual and geometric fidelity, editing controllability, and cross-modality consistency. Furthermore, generating rare-category vehicle data with MultiEditor substantially enhances the detection accuracy of perception models on underrepresented classes.

多模态编辑3D高斯自动驾驶可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。