FaceX通过区域级激活聚合,可视化人脸属性模型的全局决策依据。
FaceX: Understanding Face Attribute Classifiers through Summary Model Explanations
- 用19个面部区域的激活聚合,生成模型决策的总结性解释。
- 在4个数据集上验证,能有效识别模型中的性别、种族等偏见。
- 适合关注人脸模型公平性的研究人员和产品开发者。
可解释人工智能(XAI)方法广泛用于识别人工智能系统中的公平性问题。然而,在人脸分析领域,现有XAI方法如像素归因法仅提供单张图像的解释,难以评估模型整体行为,需人工检查大量样本才能形成总体印象,效率低下。为此,我们提出FaceX,首个通过总结性模型解释全面理解人脸属性分类器的方法。FaceX利用所有面部图像中存在的一致区域,计算区域级模型激活的聚合结果,实现对19个预定义面部区域(如头发、耳朵、皮肤)的模型区域归因可视化。此外,它还通过可视化每个面部区域在测试基准中影响模型决策最强的图像片段,提升可解释性。在多个实验设置下(包括有意引入偏见及缓解策略),涵盖CelebA、FairFace、CelebAMask-HQ和Racial Faces in the Wild四个基准,实证表明FaceX在识别模型偏见方面具有高度有效性。
原文摘要 · Abstract (English)
EXplainable Artificial Intelligence (XAI) approaches are widely applied for identifying fairness issues in Artificial Intelligence (AI) systems. However, in the context of facial analysis, existing XAI approaches, such as pixel attribution methods, offer explanations for individual images, posing challenges in assessing the overall behavior of a model, which would require labor-intensive manual inspection of a very large number of instances and leaving to the human the task of drawing a general impression of the model behavior from the individual outputs. Addressing this limitation, we introduce FaceX, the first method that provides a comprehensive understanding of face attribute classifiers through summary model explanations. Specifically, FaceX leverages the presence of distinct regions across all facial images to compute a region-level aggregation of model activations, allowing for the visualization of the model's region attribution across 19 predefined regions of interest in facial images, such as hair, ears, or skin. Beyond spatial explanations, FaceX enhances interpretability by visualizing specific image patches with the highest impact on the model's decisions for each facial region within a test benchmark. Through extensive evaluation in various experimental setups, including scenarios with or without intentional biases and mitigation efforts on four benchmarks, namely CelebA, FairFace, CelebAMask-HQ, and Racial Faces in the Wild, FaceX demonstrates high effectiveness in identifying the models' biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。