arXiv:2412.17523cs.LGcs.AI2024-12AAAI被引 2

构建公平潜在空间,让模型解释更可信且无偏。

Constructing Fair Latent Space for Intersection of Fairness and Explainability

  • 通过解耦与重分配标签和敏感属性构建公平潜在空间。
  • 在多个公平性指标上验证,有效生成反事实解释并保证公平性。
  • 仅需训练模块,低成本适配现有生成模型,适合可解释性研究者。

随着机器学习模型应用增加,公平性研究日益重要,但公平性与可解释性交叉领域仍不充分,影响用户信任。本文提出一种新模块,构建公平潜在空间,在确保公平的同时实现忠实解释。该模块通过解耦并重分配标签与敏感属性,生成各类信息的反事实解释。模块可附加于预训练生成模型,将有偏潜在空间转化为公平空间,且仅需训练模块,节省时间和成本。我们通过多种公平性指标验证其有效性,结果表明该方法能有效解释偏差决策并提供公平性保障。

原文摘要 · Abstract (English)

As the use of machine learning models has increased, numerous studies have aimed to enhance fairness. However, research on the intersection of fairness and explainability remains insufficient, leading to potential issues in gaining the trust of actual users. Here, we propose a novel module that constructs a fair latent space, enabling faithful explanation while ensuring fairness. The fair latent space is constructed by disentangling and redistributing labels and sensitive attributes, allowing the generation of counterfactual explanations for each type of information. Our module is attached to a pretrained generative model, transforming its biased latent space into a fair latent space. Additionally, since only the module needs to be trained, there are advantages in terms of time and cost savings, without the need to train the entire generative model. We validate the fair latent space with various fairness metrics and demonstrate that our approach can effectively provide explanations for biased decisions and assurances of fairness.

公平性可解释性潜在空间生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。