arXiv:2605.29471cs.CV2026-05

让多车视角的街景图像生成更真实一致,提升自动驾驶协同感知能力。

V2VCrafter: Consistent Street-View Image Generation Across Vehicles

论文配图:V2VCrafter: Consistent Street-View Image Generation Across Vehicles
图 1 · 摘自论文原文
  • 基于单车扩散模型,分步引导多车图像生成。
  • 跨车一致性提升37.2%(在真实数据集上)。
  • 适合自动驾驶数据增强与多智能体感知研究者。

车联网与自动驾驶系统依赖车对车(V2V)通信实现多智能体协同感知,但受限于标注真实世界V2V数据集稀缺及跨驾驶场景泛化能力不足。图像生成可作为数据增强的有效方案,但现有单车多视角生成框架在多智能体场景下面临两大挑战:(1)扩展的学习目标降低生成质量;(2)动态的智能体间差异阻碍对共同观测物体(如颜色、类别)物理属性的一致性建模。为此,我们提出V2VCrafter,首个面向跨车可控且真实的多视角驾驶图像生成框架。为实现高效学习,我们基于单智能体主干网络构建渐进式多智能体扩散模型,利用邻近智能体的潜在状态逐步引导单到多智能体生成。为解决跨车不一致性问题,进一步提出跨智能体注意力模块,通过协作视图图与可学习的共同观测物体表征,建模动态跨车相机视角关系。在真实世界V2X-Real数据集上的实验表明,V2VCrafter生成高质量、可控且跨车一致的街景图像,显著提升下游协同3D物体检测任务性能。

原文摘要 · Abstract (English)

Connected and autonomous driving (CAD) systems leverage vehicle-to-vehicle (V2V) communication for multi-agent collaborative perception, yet remain constrained by scarce annotated real-world V2V datasets and limited generalization across diverse driving conditions. While image generation offers a feasible solution for data augmentation, existing single-vehicle multi-view generation frameworks face two key challenges in multi-agent settings: (1) the expanded learning objective degrades generation quality, and (2) dynamic inter-agent variation hinders consistency modeling for physical attributes (e.g., color, category) of jointly observed objects. To bridge this gap, we propose V2VCrafter, the first framework for generating controllable and realistic multi-view driving images across vehicles. For effective learning, we develop a progressive multi-agent diffusion model based on a single-agent backbone, using neighboring agents' latent states to progressively guide single-to-multi-agent generation. To address cross-vehicle inconsistency, we further propose a cross-agent attention module that leverages a collaboration view graph and learnable jointly observed object representations to model dynamic cross-vehicle camera view relationships. Experiments on real-world V2X-Real dataset show that V2VCrafter generates high-fidelity, controllable, and consistent street views across vehicles, thereby effectively enhancing downstream collaborative 3D object detection tasks.

图像生成自动驾驶多智能体一致性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。