arXiv:2507.00153cs.CV2025-07

用扩散模型生成雪地图像,提升自动驾驶感知在极端环境的泛化能力

Diffusion-Based Image Augmentation for Semantic Segmentation in Outdoor Robotics

  • 基于扩散模型生成雪地场景图像,增强训练数据多样性
  • 通过语义分割过滤幻觉内容,确保生成图像真实可用
  • 方法可扩展至沙地、火山等地形,适用于多种户外机器人部署

基于学习的感知算法在分布外和代表性不足的环境中表现下降。户外机器人尤其易受动态光照、季节变化和天气影响,导致实际场景与训练数据不匹配。本文聚焦自动驾驶车辆在积雪环境中的部署问题,提出一种基于扩散模型的图像增强方法,使训练数据更贴近真实部署环境。该方法依赖互联网规模数据训练的视觉基础模型,可控制地面语义分布,并对模型进行针对性微调。采用开放词汇语义分割模型筛选含幻觉的生成候选,确保数据质量。我们认为该方法可推广至沙地、火山等其他特殊地形。

原文摘要 · Abstract (English)

The performance of leaning-based perception algorithms suffer when deployed in out-of-distribution and underrepresented environments. Outdoor robots are particularly susceptible to rapid changes in visual scene appearance due to dynamic lighting, seasonality and weather effects that lead to scenes underrepresented in the training data of the learning-based perception system. In this conceptual paper, we focus on preparing our autonomous vehicle for deployment in snow-filled environments. We propose a novel method for diffusion-based image augmentation to more closely represent the deployment environment in our training data. Diffusion-based image augmentations rely on the public availability of vision foundation models learned on internet-scale datasets. The diffusion-based image augmentations allow us to take control over the semantic distribution of the ground surfaces in the training data and to fine-tune our model for its deployment environment. We employ open vocabulary semantic segmentation models to filter out augmentation candidates that contain hallucinations. We believe that diffusion-based image augmentations can be extended to many other environments apart from snow surfaces, like sandy environments and volcanic terrains.

图像增强扩散模型语义分割自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。