用反事实分析揭示视觉模型的偏见,让AI更公平可信。
Understanding and evaluating computer vision models through the lens of counterfactuals
- 通过改变输入语义属性,观察模型响应变化来探查偏见
- 发现模型依赖背景等无关线索,且在生成中存在多重身份偏见
- 提供可落地的工具链,适合关注AI公平性的研究者与开发者
反事实推理——通过改变输入并观察模型行为变化来回答‘如果……会怎样’——已成为可解释性与公平性AI的核心。本文构建了基于反事实的框架,用于解释、审计和缓解视觉分类器与生成模型中的偏见。通过系统性地修改语义相关属性而固定其他因素,该方法揭示了虚假相关性,探测因果依赖关系,并提升系统鲁棒性。第一部分针对分类模型:CAVLI融合归因(LIME)与概念级分析(TCAV),量化决策对人类可理解概念的依赖程度;结合局部热图与概念依赖评分,揭示模型是否依赖背景等无关线索。ASAC引入对抗性反事实,扰动受保护属性的同时保持语义不变,通过课程学习微调有偏模型,在提升公平性与准确率的同时避免刻板印象。第二部分面向生成式文本到图像(TTI)模型:TIBET提供可扩展的评估流程,通过替换身份相关词检测提示敏感偏见,实现因果审计;BiasConnect构建因果图诊断交叉偏见;InterMit则提出一种模块化、无需训练的算法,利用因果敏感度与用户定义的公平目标缓解交叉偏见。这些工作表明,反事实是统一解释性、公平性与因果性的核心视角,为社会负责任的偏见评估与缓解提供了原则性、可扩展的方法。
原文摘要 · Abstract (English)
Counterfactual reasoning -- the practice of asking ``what if'' by varying inputs and observing changes in model behavior -- has become central to interpretable and fair AI. This thesis develops frameworks that use counterfactuals to explain, audit, and mitigate bias in vision classifiers and generative models. By systematically altering semantically meaningful attributes while holding others fixed, these methods uncover spurious correlations, probe causal dependencies, and help build more robust systems. The first part addresses vision classifiers. CAVLI integrates attribution (LIME) with concept-level analysis (TCAV) to quantify how strongly decisions rely on human-interpretable concepts. With localized heatmaps and a Concept Dependency Score, CAVLI shows when models depend on irrelevant cues like backgrounds. Extending this, ASAC introduces adversarial counterfactuals that perturb protected attributes while preserving semantics. Through curriculum learning, ASAC fine-tunes biased models for improved fairness and accuracy while avoiding stereotype-laden artifacts. The second part targets generative Text-to-Image (TTI) models. TIBET provides a scalable pipeline for evaluating prompt-sensitive biases by varying identity-related terms, enabling causal auditing of how race, gender, and age affect image generation. To capture interactions, BiasConnect builds causal graphs diagnosing intersectional biases. Finally, InterMit offers a modular, training-free algorithm that mitigates intersectional bias via causal sensitivity scores and user-defined fairness goals. Together, these contributions show counterfactuals as a unifying lens for interpretability, fairness, and causality in both discriminative and generative models, establishing principled, scalable methods for socially responsible bias evaluation and mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。