用扩散模型实现跨模态物体无缝插入,提升自动驾驶测试真实性。
MObI: Multimodal Object Inpainting Using Diffusion Models
- 以3D框为条件,控制物体位置与大小,实现精准定位。
- 单张图像可同时生成相机与激光雷达的逼真补全结果。
- 适合用于感知模型测试,提升合成数据可控性与真实性。
安全关键应用如自动驾驶需要大量多模态数据进行严格测试。由于真实世界数据获取成本高、难度大,基于合成数据的方法日益重要,但要求具备高度真实感与可控性。本文提出MObI框架,利用扩散模型实现跨感知模态的多模态物体修复,同时在相机与激光雷达数据上验证。仅需一张参考RGB图像,MObI即可根据3D边界框指定位置,将物体无缝插入现有多模态场景中,保持语义一致性和多模态一致性。不同于依赖编辑掩码的传统方法,本方法通过3D边界框条件实现物体精确定位与真实缩放。结果表明,该方法能灵活插入新物体,显著提升感知模型测试中合成数据的质量与实用性。
原文摘要 · Abstract (English)
Safety-critical applications, such as autonomous driving, require extensive multimodal data for rigorous testing. Methods based on synthetic data are gaining prominence due to the cost and complexity of gathering real-world data but require a high degree of realism and controllability in order to be useful. This paper introduces MObI, a novel framework for Multimodal Object Inpainting that leverages a diffusion model to create realistic and controllable object inpaintings across perceptual modalities, demonstrated for both camera and lidar simultaneously. Using a single reference RGB image, MObI enables objects to be seamlessly inserted into existing multimodal scenes at a 3D location specified by a bounding box, while maintaining semantic consistency and multimodal coherence. Unlike traditional inpainting methods that rely solely on edit masks, our 3D bounding box conditioning gives objects accurate spatial positioning and realistic scaling. As a result, our approach can be used to insert novel objects flexibly into multimodal scenes, providing significant advantages for testing perception models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。