通过融合残差特征生成多样化图像,提升度量学习的类内多样性。
BLENDER: Blended Text Embeddings and Diffusion Residuals for Intra-Class Image Synthesis in Deep Metric Learning
- 利用去噪残差的并集与交集操作,可控合成类内属性组合。
- 在CUB-200上召回率提升3.7%,Cars-196上提升1.8%。
- 适合需要增强类内多样性的度量学习数据增强场景。
深度生成模型(DGM)的发展使得高质量合成数据成为可能。当用于增强深度度量学习(DML)中的真实数据时,这些合成样本能提升类内多样性,并改善下游DML任务性能。我们提出BLenDeR,一种基于去噪残差集合论操作的扩散采样方法,可可控地增强DML中的类内多样性。并集操作鼓励多个提示中任意出现的属性,交集则通过主成分代理提取共有方向。该方法实现每个类别内多样化属性组合的可控合成,解决了现有生成方法的关键局限。在标准DML基准上的实验表明,BLenDeR在多个数据集和骨干网络下均优于当前最优基线。具体而言,在标准设置下,相比现有最优方法,于CUB-200上取得3.7%的Recall@1提升,于Cars-196上提升1.8%。
原文摘要 · Abstract (English)
The rise of Deep Generative Models (DGM) has enabled the generation of high-quality synthetic data. When used to augment authentic data in Deep Metric Learning (DML), these synthetic samples enhance intra-class diversity and improve the performance of downstream DML tasks. We introduce BLenDeR, a diffusion sampling method designed to increase intra-class diversity for DML in a controllable way by leveraging set-theory inspired union and intersection operations on denoising residuals. The union operation encourages any attribute present across multiple prompts, while the intersection extracts the common direction through a principal component surrogate. These operations enable controlled synthesis of diverse attribute combinations within each class, addressing key limitations of existing generative approaches. Experiments on standard DML benchmarks demonstrate that BLenDeR consistently outperforms state-of-the-art baselines across multiple datasets and backbones. Specifically, BLenDeR achieves 3.7% increase in Recall@1 on CUB-200 and a 1.8% increase on Cars-196, compared to state-of-the-art baselines under standard experimental settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。