arXiv:2507.02217cs.CVcs.AI2025-07

对比两种合成数据生成方法,发现布局条件更优。

Understanding Trade offs When Conditioning Synthetic Data

  • 用提示词或布局图控制合成图像生成
  • 布局条件在多样化场景下提升精度34%~177%
  • 适合小样本视觉检测与机器人应用

从少量图像中训练鲁棒的目标检测器是工业视觉系统中的关键挑战,因高质量数据收集需数月。合成数据成为数据高效视觉检测和抓取机器人的主要解决方案。当前流程依赖Blender、Unreal等3D引擎,虽有精细控制但渲染小数据集需数周,且仿真与现实差距大。扩散模型可分钟级生成高质量图像,但低数据场景下的精准控制仍困难。尽管已有多种适配器扩展扩散模型的文本提示能力,不同条件策略对合成数据质量的影响尚不明确。我们研究了来自四个标准目标检测基准的80种视觉概念,对比提示词条件与布局条件两种策略。当条件线索较窄时,提示词条件生成数据质量更高;随多样性增加,布局条件表现更优。当布局线索覆盖完整训练分布时,合成数据使平均精度提升34%(最高达177%),优于仅使用真实数据。

原文摘要 · Abstract (English)

Learning robust object detectors from only a handful of images is a critical challenge in industrial vision systems, where collecting high quality training data can take months. Synthetic data has emerged as a key solution for data efficient visual inspection and pick and place robotics. Current pipelines rely on 3D engines such as Blender or Unreal, which offer fine control but still require weeks to render a small dataset, and the resulting images often suffer from a large gap between simulation and reality. Diffusion models promise a step change because they can generate high quality images in minutes, yet precise control, especially in low data regimes, remains difficult. Although many adapters now extend diffusion beyond plain text prompts, the effect of different conditioning schemes on synthetic data quality is poorly understood. We study eighty diverse visual concepts drawn from four standard object detection benchmarks and compare two conditioning strategies: prompt based and layout based. When the set of conditioning cues is narrow, prompt conditioning yields higher quality synthetic data; as diversity grows, layout conditioning becomes superior. When layout cues match the full training distribution, synthetic data raises mean average precision by an average of thirty four percent and by as much as one hundred seventy seven percent compared with using real data alone.

合成数据扩散模型目标检测布局条件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。