用扩散模型同步生成车载多传感器数据,提升真实感与跨模态一致性。
X-Drive: Cross-modality consistent multi-sensor data synthesis for driving scenarios
- 双分支潜空间扩散架构,通过跨模态局部区域条件建模联合分布。
- 基于极线设计的交叉条件模块,缓解去噪过程中的空间模糊问题。
- 支持文本、框、图像、点云等多级输入,可控生成且保持模态间一致。
近期研究利用扩散模型在自动驾驶场景中合成激光雷达点云或相机图像数据。尽管其在建模单模态数据边缘分布方面取得成功,但对多模态间相互依赖关系的探索仍不足。为此,本文提出X-DRIVE框架,通过双分支潜空间扩散模型架构,建模点云与多视角图像的联合分布。鉴于两模态几何空间差异,X-DRIVE将每种模态的生成条件设置为另一模态的对应局部区域,确保更好对齐与真实感。为进一步处理去噪过程中的空间模糊性,设计基于极线的跨模态条件模块,自适应学习跨模态局部对应关系。此外,X-DRIVE支持多层级输入条件(如文本、边界框、图像、点云),实现可控生成,同时保证可靠跨模态一致性。大量实验表明,该方法在点云与多视角图像合成上均能生成高保真结果,且严格遵循输入条件。
原文摘要 · Abstract (English)
Recent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an under-exploration in the mutual reliance between different modalities to describe complex driving scenes. To fill in this gap, we propose a novel framework, X-DRIVE, to model the joint distribution of point clouds and multi-view images via a dual-branch latent diffusion model architecture. Considering the distinct geometrical spaces of the two modalities, X-DRIVE conditions the synthesis of each modality on the corresponding local regions from the other modality, ensuring better alignment and realism. To further handle the spatial ambiguity during denoising, we design the cross-modality condition module based on epipolar lines to adaptively learn the cross-modality local correspondence. Besides, X-DRIVE allows for controllable generation through multi-level input conditions, including text, bounding box, image, and point clouds. Extensive results demonstrate the high-fidelity synthetic results of X-DRIVE for both point clouds and multi-view images, adhering to input conditions while ensuring reliable cross-modality consistency. Our code will be made publicly available at https://github.com/yichen928/X-Drive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。