越好看的图,越不适合当训练数据,新模型反而更差。
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
- 用最新扩散模型生成大量合成图像做训练
- 新模型生成的数据让分类器在真实数据上准确率下降
- 图像太美导致分布单一,反而丢失真实数据多样性
近期文本到图像(T2I)扩散模型生成的图像视觉效果惊艳且能精准遵循提示词。但它们能否作为可靠的合成视觉数据生成器?本文重新审视了合成数据作为真实训练集替代方案的潜力,发现一个令人意外的性能退化现象。我们使用2022至2025年间发布的顶尖T2I模型生成大规模合成数据集,仅用这些合成数据训练标准分类器,并在真实测试集上评估其表现。尽管视觉保真度和提示词遵循能力持续提升,分类准确率却随模型版本更新而不断下降。分析揭示出隐藏趋势:这些模型逐渐收敛至狭窄的、以美学为中心的分布,严重削弱了数据多样性和对真实数据分布的覆盖。研究挑战了视觉领域一个日益普遍的假设——生成图像越逼真,其作为训练数据的价值就越高。因此,亟需重新评估现代T2I模型作为可靠训练数据生成器的能力。
原文摘要 · Abstract (English)
Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as synthetic vision data generators? In this work, we revisit the promise of synthetic data as a scalable substitute for real training sets and uncover a surprising performance regression. We generate large-scale synthetic datasets using state-of-the-art T2I models released between 2022 and 2025, train standard classifiers solely on this synthetic data, and evaluate them on real test data. Despite observable advances in visual fidelity and prompt adherence, classification accuracy on real test data consistently declines with newer T2I models as training data generators. Our analysis reveals a hidden trend: These models collapse to a narrow, aesthetic-centric distribution that undermines diversity and real data distribution coverage. Overall, our findings challenge a growing assumption in vision research, namely that progress in generative realism implies progress in data realism. We thus highlight an urgent need to rethink the capabilities of modern T2I models as reliable training data generators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。