用高质量合成图像提升图像分割效果,关键在场景密集和实例精细。
What Makes Synthetic Data Effective in Image Segmentation

- 通过分析扩散模型生成的合成图像,发现密集场景和精细实例更有效。
- 在城市街景、COCO、ADE20K上,新方法显著提升分割精度。
- 框架通用且可扩展,适配多种模型,适合做视觉分割研究者使用。
随着大规模生成模型的快速发展,合成数据成为视觉理解的有力解决方案。尽管现代扩散模型能生成逼真的图像,其在复杂视觉分割任务中的潜力仍待深入探索。本文系统分析了先进扩散模型生成的合成图像,揭示其有效性关键因素:场景密集布局与精细实例保真度带来更具区分性的空间表征。基于此,我们提出 SENSE 框架,利用灵活可扩展的合成数据显著提升分割性能。SENSE 兼容多种模型架构(如 DPT、Mask2Former),对不同参数规模模型均具良好扩展性。在 Cityscapes、COCO 与 ADE20K 上的大量实验验证了方法的有效性与泛化能力。代码已开源。
原文摘要 · Abstract (English)
Driven by rapid advances in large-scale generative models, synthetic data has emerged as a promising solution for visual understanding. While modern diffusion models achieve remarkable photorealistic image synthesis, their potential in complex visual segmentation tasks remains underexplored. In this work, we conduct a systematic analysis of synthetic images from state-of-the-art diffusion models to uncover the factors governing their utility. In particular, synthetic images characterized by dense scene composition and fine instance fidelity demonstrate distinctive benefits, yielding significantly more discriminative spatial representations. Building on these insights, we propose SENSE, a unified framework that leverages flexible and scalable synthetic data to substantially enhance segmentation performance. Notably, SENSE is model-agnostic, compatible with diverse architectures (e.g., DPT and Mask2Former), and scales effectively across models with varying parameter capacities. Extensive experiments on Cityscapes, COCO, and ADE20K validate the effectiveness and generalization capability of our approach. Code is available at https://github.com/zhang0jhon/SENSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。