arXiv:2410.12075cs.CVcs.AI2024-10被引 1

用大模型生成多样天气驾驶图像,提升分割模型泛化能力

WeatherDG: LLM-assisted Diffusion Model for Procedural Weather Generation in Domain-Generalized Semantic Segmentation

  • 结合扩散模型与大语言模型,自动构建丰富天气场景描述
  • 在Cityscapes→ACDC任务上使基线模型mIoU提升13.9%
  • 无需修改分割模型,可通用增强各类现成模型的鲁棒性

本文提出WeatherDG,一种基于稳定扩散(Stable Diffusion, SD)与大语言模型(LLM)协同的程序化天气生成方法。首先微调SD以匹配真实驾驶场景的内容与布局;随后设计基于LLM的程序化提示生成机制,丰富场景描述,驱动SD自动生成更多样、更精细的天气图像。此外,引入均衡生成策略,确保在各种天气条件下均能高质量生成尾部类别(如骑行者、摩托车)目标。该分割模型无关的方法通过合成数据增强,显著提升现有模型在目标域上的泛化性能。在三个挑战性数据集上的实验表明,该方法可显著提升多种先进分割模型的表现。特别地,在Cityscapes→ACDC设置下,相比基线HRDA模型,mIoU提升13.9%。

原文摘要 · Abstract (English)

In this work, we propose a novel approach, namely WeatherDG, that can generate realistic, weather-diverse, and driving-screen images based on the cooperation of two foundation models, i.e, Stable Diffusion (SD) and Large Language Model (LLM). Specifically, we first fine-tune the SD with source data, aligning the content and layout of generated samples with real-world driving scenarios. Then, we propose a procedural prompt generation method based on LLM, which can enrich scenario descriptions and help SD automatically generate more diverse, detailed images. In addition, we introduce a balanced generation strategy, which encourages the SD to generate high-quality objects of tailed classes under various weather conditions, such as riders and motorcycles. This segmentation-model-agnostic method can improve the generalization ability of existing models by additionally adapting them with the generated synthetic data. Experiments on three challenging datasets show that our method can significantly improve the segmentation performance of different state-of-the-art models on target domains. Notably, in the setting of ''Cityscapes to ACDC'', our method improves the baseline HRDA by 13.9% in mIoU.

天气生成扩散模型领域泛化语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。