通过自适应生成与质量引导监督,提升对比学习的图像对质量与效果。
GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
- 动态调节生成参数,融合多源策略生成多样且语义一致的图像视图。
- 根据语义对齐度与多样性重加权,优先训练高质量样本对。
- 在图像和图文任务中显著提升性能,适合视觉与跨模态学习研究者。
对比学习的成功依赖于高质量正样本对的构建与利用。然而,现有方法在构建与学习两方面均存在关键局限:构建方面,手工与生成增强常导致多样性不足且易引发语义扭曲;学习方面,缺乏质量评估机制,导致所有样本对被同等对待,造成次优监督。为此,我们提出GenView++,一个统一框架,通过两项协同创新解决上述问题。首先,引入多源自适应视图生成机制,通过动态调节图像、文本及图文联合策略下的生成参数,合成多样化且语义一致的视图。其次,设计质量驱动的对比学习机制,评估每对样本的语义对齐度与多样性,动态重加权其训练贡献,优先优化高质量对,抑制冗余或错位对。大量实验表明,GenView++在视觉与视觉-语言任务中均具显著有效性:在图像表示学习中,使MoCov2在ImageNet线性分类上提升+2.5%;在视觉-语言学习中,在十个数据集上平均零样本分类准确率较CLIP提升+12.31%,较SLIP提升+5.31%,并在Flickr30k文本检索任务中,R@5指标提升+3.2%。
原文摘要 · Abstract (English)
The success of contrastive learning depends on the construction and utilization of high-quality positive pairs. However, current methods face critical limitations on two fronts: on the construction side, both handcrafted and generative augmentations often suffer from limited diversity and risk semantic corruption; on the learning side, the absence of a quality assessment mechanism leads to suboptimal supervision where all pairs are treated equally. To tackle these challenges, we propose GenView++, a unified framework that addresses both fronts by introducing two synergistic innovations. To improve pair construction, GenView++ introduces a multi-source adaptive view generation mechanism to synthesize diverse yet semantically coherent views by dynamically modulating generative parameters across image-conditioned, text-conditioned, and image-text-conditioned strategies. Second, a quality-driven contrastive learning mechanism assesses each pair's semantic alignment and diversity to dynamically reweight their training contribution, prioritizing high-quality pairs while suppressing redundant or misaligned pairs. Extensive experiments demonstrate the effectiveness of GenView++ across both vision and vision-language tasks. For vision representation learning, it improves MoCov2 by +2.5% on ImageNet linear classification. For vision-language learning, it raises the average zero-shot classification accuracy by +12.31% over CLIP and +5.31% over SLIP across ten datasets, and further improves Flickr30k text retrieval R@5 by +3.2%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。