arXiv:2510.11567cs.CVcs.GR2025-10

用扩散模型快速生成与真实城市图像对齐的语义分割数据,提升训练效果。

A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation

  • 基于伪标签微调扩散模型,将合成图转为真实感目标域图像。
  • 在5个合成数据集上实现最高+8.0% mIoU提升,媲美耗时数月的设计。
  • 适合需要快速构建高质量训练数据的城市视觉研究者。

合成数据广泛用于城市场景识别模型训练,但即使高度逼真的渲染也与真实图像存在显著差异。这种差距在适配特定目标域(如Cityscapes)时尤为明显,因建筑风格、植被、物体外观和相机特性差异导致下游性能下降。若通过更精细的3D建模弥补差距,将大幅增加成本,违背低成本标注数据的初衷。为此,我们提出一种新框架,仅使用不完美的伪标签即可将现成扩散模型适配至目标域。训练完成后,该模型能从任意合成数据集的语义图生成高保真、与目标域对齐的图像,包括仅用数小时构建的低努力数据源。方法包含筛选劣质生成、修正图像-标签错位、统一跨数据集语义等步骤,将弱合成数据转化为可竞争的真实域训练集。在五个合成数据集和两个真实目标数据集上的实验表明,分割精度提升最高达+8.0% mIoU,优于当前最先进的图像转换方法,使快速构建的合成数据达到需长期人工设计的高质量合成数据的效果。本工作揭示了一种高效协作范式:快速语义原型结合生成模型,可实现城市场景理解中可扩展、高质量训练数据的创建。

原文摘要 · Abstract (English)

Synthetic datasets are widely used for training urban scene recognition models, but even highly realistic renderings show a noticeable gap to real imagery. This gap is particularly pronounced when adapting to a specific target domain, such as Cityscapes, where differences in architecture, vegetation, object appearance, and camera characteristics limit downstream performance. Closing this gap with more detailed 3D modelling would require expensive asset and scene design, defeating the purpose of low-cost labelled data. To address this, we present a new framework that adapts an off-the-shelf diffusion model to a target domain using only imperfect pseudo-labels. Once trained, it generates high-fidelity, target-aligned images from semantic maps of any synthetic dataset, including low-effort sources created in hours rather than months. The method filters suboptimal generations, rectifies image-label misalignments, and standardises semantics across datasets, transforming weak synthetic data into competitive real-domain training sets. Experiments on five synthetic datasets and two real target datasets show segmentation gains of up to +8.0%pt. mIoU over state-of-the-art translation methods, making rapidly constructed synthetic datasets as effective as high-effort, time-intensive synthetic datasets requiring extensive manual design. This work highlights a valuable collaborative paradigm where fast semantic prototyping, combined with generative models, enables scalable, high-quality training data creation for urban scene understanding.

语义分割生成模型数据合成城市感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。