arXiv:2507.18763cs.CVcs.RO2025-07被引 2

用扩散模型从单目摄像头预测自动驾驶可行驶路径走廊。

Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving

  • 基于未来车辆轨迹和前视图像自监督生成行驶走廊样本。
  • 提出ContourDiff模型,通过轮廓点去噪实现结构化路径预测。
  • 在nuScenes和CARLA上验证,能准确预测多模态安全行驶区域。

可行驶自由空间预测是自动驾驶中的核心问题。现有方法通常将整个无障碍道路区域视为自由空间,而本文关注的是其中可导航的驾驶走廊。传统方法依赖鸟瞰图(BEV)表示,难以获取。为此,本文将走廊预测视为纯图像感知任务,仅使用单目摄像头输入。由于图像中缺乏对应走廊标注,我们提出一种自监督方法,利用未来车辆轨迹与前视图像生成自由空间样本,使视觉走廊估计依赖于车辆自身轨迹。随后,采用扩散过程建模图像中此类路段的分布。但现有二值掩码表示存在局限,因此提出ContourDiff——一种基于轮廓点去噪的扩散架构,实现结构化且可解释的自由空间预测。在nuScenes和CARLA数据集上进行定性与定量评估,结果表明该方法能有效预测图像中的安全多模态可行驶走廊。

原文摘要 · Abstract (English)

Drivable Free-space prediction is a fundamental and crucial problem in autonomous driving. Recent works have addressed the problem by representing the entire non-obstacle road regions as the free-space. In contrast our aim is to estimate the driving corridors that are a navigable subset of the entire road region. Unfortunately, existing corridor estimation methods directly assume a BEV-centric representation, which is hard to obtain. In contrast, we frame drivable free-space corridor prediction as a pure image perception task, using only monocular camera input. However such a formulation poses several challenges as one doesn't have the corresponding data for such free-space corridor segments in the image. Consequently, we develop a novel self-supervised approach for free-space sample generation by leveraging future ego trajectories and front-view camera images, making the process of visual corridor estimation dependent on the ego trajectory. We then employ a diffusion process to model the distribution of such segments in the image. However, the existing binary mask-based representation for a segment poses many limitations. Therefore, we introduce ContourDiff, a specialized diffusion-based architecture that denoises over contour points rather than relying on binary mask representations, enabling structured and interpretable free-space predictions. We evaluate our approach qualitatively and quantitatively on both nuScenes and CARLA, demonstrating its effectiveness in accurately predicting safe multimodal navigable corridors in the image.

自动驾驶扩散模型路径规划图像感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。