针对图像分类中的交叉偏差,提出可量化分析与增强的解决方案。
Data-Driven Analysis of Intersectional Bias in Image Classification: A Framework with Bias-Weighted Augmentation
- 用公平性指标+可解释工具定位模型偏差模式。
- 新数据增强法使少数群体准确率提升24个百分点。
- 适合关注算法公平性的研究者与开发者。
在数据不平衡的图像分类任务中,机器学习模型常表现出由物体类别与环境条件等多重属性交互引发的交叉偏差。本文提出一种数据驱动框架,通过结合定量公平性度量与可解释性工具,系统识别模型预测中的偏差模式。基于此分析,提出偏置加权增强(BWA)策略,根据子群体分布统计自适应调整变换强度。在 Open Images V7 数据集上,针对五类物体的实验表明,该方法使少数类别-环境交集的准确率最高提升24个百分点,公平性指标差异降低35%。多次独立运行的统计分析验证了改进效果的显著性(p < 0.05)。该方法为图像分类系统中的交叉偏差分析与缓解提供了可复现的路径。
原文摘要 · Abstract (English)
Machine learning models trained on imbalanced datasets often exhibit intersectional biases-systematic errors arising from the interaction of multiple attributes such as object class and environmental conditions. This paper presents a data-driven framework for analyzing and mitigating such biases in image classification. We introduce the Intersectional Fairness Evaluation Framework (IFEF), which combines quantitative fairness metrics with interpretability tools to systematically identify bias patterns in model predictions. Building on this analysis, we propose Bias-Weighted Augmentation (BWA), a novel data augmentation strategy that adapts transformation intensities based on subgroup distribution statistics. Experiments on the Open Images V7 dataset with five object classes demonstrate that BWA improves accuracy for underrepresented class-environment intersections by up to 24 percentage points while reducing fairness metric disparities by 35%. Statistical analysis across multiple independent runs confirms the significance of improvements (p < 0.05). Our methodology provides a replicable approach for analyzing and addressing intersectional biases in image classification systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。