改进概念瓶颈模型,让图像分类更公平且可解释。
Mitigating Bias in Concept Bottleneck Models for Fair and Interpretable Image Classification
- 用top-k筛选减少概念信息泄露,提升公平性。
- 在ImSitu数据集上显著降低性别偏见,优于以往方法。
- 适合关注公平性与可解释性的计算机视觉研究者。
确保图像分类的公平性可防止模型延续和放大偏见。概念瓶颈模型(CBM)通过将图像映射到高层、人类可理解的概念,再经由稀疏单层分类器进行预测,该结构增强了可解释性,并理论上可通过隐藏敏感属性代理(如面部特征)来支持公平性。然而,现有研究表明CBM概念仍会泄露与概念语义无关的信息,早期实验仅在ImSitu数据集上实现性别偏见的微弱下降。本文提出三种偏见缓解技术:1. 使用top-k概念过滤减少信息泄露;2. 移除有偏概念;3. 采用对抗去偏。实验结果表明,所提方法在公平性与性能权衡上超越先前工作,验证了去偏CBM在实现公平且可解释图像分类上的重要进展。
原文摘要 · Abstract (English)
Ensuring fairness in image classification prevents models from perpetuating and amplifying bias. Concept bottleneck models (CBMs) map images to high-level, human-interpretable concepts before making predictions via a sparse, one-layer classifier. This structure enhances interpretability and, in theory, supports fairness by masking sensitive attribute proxies such as facial features. However, CBM concepts have been known to leak information unrelated to concept semantics and early results reveal only marginal reductions in gender bias on datasets like ImSitu. We propose three bias mitigation techniques to improve fairness in CBMs: 1. Decreasing information leakage using a top-k concept filter, 2. Removing biased concepts, and 3. Adversarial debiasing. Our results outperform prior work in terms of fairness-performance tradeoffs, indicating that our debiased CBM provides a significant step towards fair and interpretable image classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。