arXiv:2506.14418cs.CVcs.AI2025-06被引 1

揭示视觉属性复合稀疏性问题,提出基于罕见组合的采样优化方案。

Compositional Attribute Imbalance in Vision Datasets

  • 构建基于CLIP的视觉属性词典,自动评估图像属性。
  • 发现复合属性稀疏性显著影响模型性能,尤其在长尾分布下。
  • 结合数据增强提升罕见属性表征能力,适合长尾分类任务研究者。

视觉属性不平衡是图像分类中常见但未被充分研究的问题,严重影响模型性能与泛化能力。本文首次定义图像的一级与二级属性,并提出基于CLIP的框架构建视觉属性词典,实现图像属性的自动评估。通过系统分析单一属性与复合属性不平衡,揭示了属性稀有性对模型表现的影响机制。为此,我们提出根据样本复合属性的稀有程度调整采样概率,并集成CutMix、Fmix、SaliencyMix等数据增强技术,以增强模型对罕见属性的表征能力。在多个基准数据集上的大量实验表明,该方法能有效缓解属性不平衡问题,提升深度神经网络的鲁棒性与公平性。本研究强调建模视觉属性分布的重要性,为长尾图像分类提供了可扩展解决方案。

原文摘要 · Abstract (English)

Visual attribute imbalance is a common yet underexplored issue in image classification, significantly impacting model performance and generalization. In this work, we first define the first-level and second-level attributes of images and then introduce a CLIP-based framework to construct a visual attribute dictionary, enabling automatic evaluation of image attributes. By systematically analyzing both single-attribute imbalance and compositional attribute imbalance, we reveal how the rarity of attributes affects model performance. To tackle these challenges, we propose adjusting the sampling probability of samples based on the rarity of their compositional attributes. This strategy is further integrated with various data augmentation techniques (such as CutMix, Fmix, and SaliencyMix) to enhance the model's ability to represent rare attributes. Extensive experiments on benchmark datasets demonstrate that our method effectively mitigates attribute imbalance, thereby improving the robustness and fairness of deep neural networks. Our research highlights the importance of modeling visual attribute distributions and provides a scalable solution for long-tail image classification tasks.

长尾分类属性平衡CLIP数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。