让生成模型主动造出难样本,提升训练效果且省时。
Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
- 用学习难度作条件信号,控制生成样本的难易程度。
- 仅用10%合成数据就超越现有最佳方法,节省63.4小时算力。
- 能可视化特定类别的难点特征,适合数据质量分析。
生成模型已成为计算机视觉任务中合成训练数据的强大工具。当前方法仅关注生成图像与目标数据分布的一致性,因此只能捕捉真实数据中的常见特征,主要生成已被模型充分学习的‘易样本’,而对性能提升至关重要的罕见‘难样本’却难以有效生成。这导致需大量合成数据才能获得明显性能提升,且增益有限。为此,我们提出一种新方法,在生成过程中同时实现领域对齐与学习难度控制,可高效生成有价值的‘难样本’,显著提升目标任务表现。该方法通过将学习难度作为额外条件信号,并设计专用编码器结构与训练-生成策略实现。多数据集实验表明,本方法在更低生成成本下取得更高性能:仅使用10%额外合成数据即可达到最优效果,相比ImageNet上先前最先进方法节省63.4 GPU小时生成时间。此外,该方法还能提供类别特异性难样本的可视化,可用于数据集分析。
原文摘要 · Abstract (English)
Generative models have become a powerful tool for synthesizing training data in computer vision tasks. Current approaches solely focus on aligning generated images with the target dataset distribution. As a result, they capture only the common features in the real dataset and mostly generate 'easy samples', which are already well learned by models trained on real data. In contrast, those rare 'hard samples', with atypical features but crucial for enhancing performance, cannot be effectively generated. Consequently, these approaches must synthesize large volumes of data to yield appreciable performance gains, yet the improvement remains limited. To overcome this limitation, we present a novel method that can learn to control the learning difficulty of samples during generation while also achieving domain alignment. Thus, it can efficiently generate valuable 'hard samples' that yield significant performance improvements for target tasks. This is achieved by incorporating learning difficulty as an additional conditioning signal in generative models, together with a designed encoder structure and training-generation strategy. Experimental results across multiple datasets show that our method can achieve higher performance with lower generation cost. Specifically, we obtain the best performance with only 10% additional synthetic data, saving 63.4 GPU hours of generation time compared to the previous SOTA on ImageNet. Moreover, our method provides insightful visualizations of category-specific hard factors, serving as a tool for analyzing datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。