arXiv:2503.00234cs.LGcs.AI2025-03

用注意力图分析模型偏见,发现去偏与去伪影可互相促进。

Investigating the Relationship Between Debiasing and Artifact Removal using Saliency Maps

  • 基于注意力图设计新评估指标,量化模型决策变化
  • 去偏方法会主动转移模型对敏感属性的关注
  • 去伪影技术可被复用于提升模型公平性

机器学习系统的广泛应用引发了对公平性和偏见的广泛关注,缓解有害偏见已成为AI发展的关键。本文研究了计算机视觉任务中去偏与消除模型伪影之间的关系。首先,提出一套基于XAI的新指标,通过分析注意力图来评估模型决策过程的变化;其次,证明有效的去偏方法会系统性地将模型关注点从受保护属性上移开;最后,表明原本用于消除伪影的技术可被有效重用于改善模型公平性。这些发现为确保公平性与消除对应于受保护属性的伪影之间存在双向关联提供了证据。

原文摘要 · Abstract (English)

The widespread adoption of machine learning systems has raised critical concerns about fairness and bias, making mitigating harmful biases essential for AI development. In this paper, we investigate the relationship between debiasing and removing artifacts in neural networks for computer vision tasks. First, we introduce a set of novel XAI-based metrics that analyze saliency maps to assess shifts in a model's decision-making process. Then, we demonstrate that successful debiasing methods systematically redirect model focus away from protected attributes. Finally, we show that techniques originally developed for artifact removal can be effectively repurposed for improving fairness. These findings provide evidence for the existence of a bidirectional connection between ensuring fairness and removing artifacts corresponding to protected attributes.

模型公平性注意力图去偏方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。