arXiv:2607.02637cs.LGcs.AI2026-07

通过筛选合成图像中的多样性样本,用更少数据达到更好效果。

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

论文配图:Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
图 1 · 摘自论文原文
  • 将图像按典型性与多样性分组,用语义一致性和去重程度评分
  • 仅用40%的合成样本就达到真实数据性能,优于现有选择方法
  • 无需重训练,适用于各种生成器,适合数据效率研究者

近期生成模型可产出高质量合成图像,为数据密集型模型提供可扩展的训练数据。现有方法通常需训练或微调生成器,或依赖轻量级后处理如提示工程,具有生成器特异性且依赖专业知识。本文提出一个互补问题:给定固定生成图像池,能否仅通过选择信息量高的子集来提升下游性能?答案是肯定的。我们发现现代生成器存在结构性偏差——过度生成每类的典型模式,而忽视类内差异。基于此,我们将真实类别划分为典型(同质,HO)和非冗余(异质,HE)两部分,对合成图像进行基于保真度-多样性准则的评分,奖励语义一致性,惩罚典型重复。该方法不依赖生成器,无需重训练。在多个基准上,其表现持续优于当前最优数据选择基线,仅使用最多40%的合成样本即可达到真实数据性能。该准则在更强任务微调生成器上也有效,分类与分割任务均有提升。因此,后生成筛选并非替代优质生成器,而是提升合成数据效用的互补机制。

原文摘要 · Abstract (English)

Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existing approaches to exploiting this potential typically involve 1) training or fine-tuning generators, or 2) using lightweight post-hoc adaptation like prompt engineering or inference-time guidance, making them generator-specific and expertise-intensive. We study a complementary question: given a fixed pool of generated images, can downstream utility be improved purely by selecting an informative subset? The answer is yes. We show that effective selection must counter a structural bias of modern generators: they tend to over-produce canonical modes of each class while underrepresenting intra-class variation. Building on this insight, we split each real class into a canonical Homogeneous (HO) subset and a non-redundant Heterogeneous (HE) subset, then score synthetic images by a fidelity-diversity criterion that rewards semantic alignment while penalizing canonical redundancy. The method is generator-agnostic and requires no retraining. Across multiple benchmarks, it consistently outperforms state-of-the-art data selection baselines and matches the real-data performance with up to 40% fewer synthetic samples. The same criterion remains effective when applied on top of stronger task-tuned generators, with gains on both classification and segmentation tasks. Post-generation selection is therefore not a substitute for better generators, but a complementary mechanism for improving the utility of synthetic data.

合成数据数据筛选生成模型多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。