用概念图诊断视觉数据偏见,提升模型泛化能力
Visual Data Diagnosis and Debiasing with Concept Graphs
- 将视觉数据建模为概念知识图谱,识别虚假共现关系
- 通过团块级概念平衡策略,显著缓解数据偏差
- 适用于需要公平性与鲁棒性的图像分类任务
深度学习模型的成功依赖于大规模、复杂数据集的构建。然而,训练过程中常会吸收数据中的固有偏见,导致预测不可靠。因此,诊断并消除数据偏见至关重要。本文提出ConBias框架,用于诊断和缓解视觉数据中的概念共现偏见。该框架将视觉数据表示为概念知识图谱,可精细分析虚假概念共现,揭示数据整体中的概念不平衡问题。此外,我们提出一种基于团块的概念平衡策略,有效缓解此类不平衡,从而提升下游任务性能。大量实验表明,基于ConBias生成的均衡概念分布进行数据增强,相比现有最优方法,在多个数据集上均提升了模型泛化能力。
原文摘要 · Abstract (English)
The widespread success of deep learning models today is owed to the curation of extensive datasets significant in size and complexity. However, such models frequently pick up inherent biases in the data during the training process, leading to unreliable predictions. Diagnosing and debiasing datasets is thus a necessity to ensure reliable model performance. In this paper, we present ConBias, a novel framework for diagnosing and mitigating Concept co-occurrence Biases in visual datasets. ConBias represents visual datasets as knowledge graphs of concepts, enabling meticulous analysis of spurious concept co-occurrences to uncover concept imbalances across the whole dataset. Moreover, we show that by employing a novel clique-based concept balancing strategy, we can mitigate these imbalances, leading to enhanced performance on downstream tasks. Extensive experiments show that data augmentation based on a balanced concept distribution augmented by Conbias improves generalization performance across multiple datasets compared to state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。