用多样化人群数据提升零样本图像分类准确率
D3G: Diverse Demographic Data Generation Increases Zero-Shot Image Classification Accuracy within Multimodal Models
- 推理时生成多样化人群图像增强模型表现
- 在多个数据集上提升准确率,最高达12.3%
- 适合关注公平性与零样本性能的研究者
图像分类是实现人类级视觉理解的关键任务。尽管如CLIP等多模态模型通过学习视觉与语言间的语义相似性取得了良好效果,但该任务仍具挑战性。低容量模型常因欠拟合而在细粒度分类中表现不佳。同时,高质量且具有丰富跨模态表征的训练数据难以获取。若数据集未均衡覆盖不同人口统计特征,模型预测将偏向高频率类别,忽略其他类别。本文聚焦这些因素如何导致零样本图像分类中的有害偏差,并提出训练无关的零样本方法D3G(Diverse Demographic Data Generation),以提升预训练多模态模型的分类准确率并减少人口统计偏差。我们以CLIP为基线模型,使用Stable Diffusion XL生成多样化人口图像,在推理阶段注入这些数据,显著提升模型性能。实验表明,多样化输入可有效改善分类准确率,最高达12.3%提升,且对不同人口特征的影响可量化分析。
原文摘要 · Abstract (English)
Image classification is a task essential for machine perception to achieve human-level image understanding. Multimodal models such as CLIP have been able to perform well on this task by learning semantic similarities across vision and language; however, despite these advances, image classification is still a challenging task. Models with low capacity often suffer from underfitting and thus underperform on fine-grained image classification. Along with this, it is important to ensure high-quality data with rich cross-modal representations of each class, which is often difficult to generate. When datasets do not enforce balanced demographics, the predictions will be biased toward the more represented class, while others will be neglected. We focus on how these issues can lead to harmful bias for zero-shot image classification, and explore how to combat these issues in demographic bias. We propose Diverse Demographic Data Generation (D3G), a training-free, zero-shot method of boosting classification accuracy while reducing demographic bias in pre-trained multimodal models. With this method, we utilize CLIP as our base multimodal model and Stable Diffusion XL as our generative model. We demonstrate that providing diverse demographic data at inference time improves performance for these models, and explore the impact of individual demographics on the resulting accuracy metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。