用分阶段引导生成更精细的合成图像,提升细粒度分类性能。
HiGFA: Hierarchical Guidance for Fine-grained Data Augmentation with Diffusion Models
- 分阶段引导:前期强文本与轮廓引导建模整体结构,后期激活细粒度分类器
- 动态调节引导强度,依据预测置信度优化生成细节
- 在多个细粒度数据集上显著提升分类准确率,适合高精度图像生成任务
生成式扩散模型在数据增强中展现潜力,但在细粒度任务中面临挑战:如何确保合成图像准确捕捉细微但关键的类别特征。标准方法如基于文本的无分类器引导(CFG)缺乏足够特异性,可能生成误导性样本,降低分类器性能。为此,我们提出分层引导的细粒度数据增强方法(HiGFA)。HiGFA利用扩散采样过程的时间动态性,在早期至中期采样阶段使用固定强度的强文本和变形轮廓引导,建立整体场景、风格与结构;在最后阶段激活专用细粒度分类器引导,并根据预测置信度动态调节所有引导信号的强度。这种分层且基于置信度的协同机制,使HiGFA能生成多样且高保真的合成图像,智能平衡全局结构与精确细节。在多个FGVC数据集上的实验表明,该方法有效提升了细粒度分类性能。
原文摘要 · Abstract (English)
Generative diffusion models show promise for data augmentation. However, applying them to fine-grained tasks presents a significant challenge: ensuring synthetic images accurately capture the subtle, category-defining features critical for high fidelity. Standard approaches, such as text-based Classifier-Free Guidance (CFG), often lack the required specificity, potentially generating misleading examples that degrade fine-grained classifier performance. To address this, we propose Hierarchically Guided Fine-grained Augmentation (HiGFA). HiGFA leverages the temporal dynamics of the diffusion sampling process. It employs strong text and transformed contour guidance with fixed strengths in the early-to-mid sampling stages to establish overall scene, style, and structure. In the final sampling stages, HiGFA activates a specialized fine-grained classifier guidance and dynamically modulates the strength of all guidance signals based on prediction confidence. This hierarchical, confidence-driven orchestration enables HiGFA to generate diverse yet faithful synthetic images by intelligently balancing global structure formation with precise detail refinement. Experiments on several FGVC datasets demonstrate the effectiveness of HiGFA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。