arXiv:2509.25662cs.AIcs.SC2025-09中稿 · AJCAI 2025被引 1

用逻辑推理法揭示AI决策中隐藏的不公平代理特征

On Explaining Proxy Discrimination and Unfairness in Individual Decisions Made by AI Systems

  • 通过反向推理识别导致不公平的非直接敏感特征
  • 引入能力属性实现跨群体公平性量化评估
  • 适用于金融信贷等高风险决策系统的公平性审计

高风险领域的人工智能系统引发对代理歧视、不公平和可解释性的担忧。现有审计方法往往无法揭示不公平的根本原因,尤其当其源于结构性偏见时。本文提出一种新框架,利用形式化反向解释来阐明个体AI决策中的代理歧视。结合背景知识,该方法识别出哪些特征作为受保护属性的不合理代理,揭示隐藏的结构性偏见。核心在于引入‘能力’这一与群体归属无关的任务相关属性,并通过映射函数将不同群体中能力相当的个体进行对齐,从而实质性地评估公平性。以德国信用数据集为例,展示了该框架在真实场景中的适用性。

原文摘要 · Abstract (English)

Artificial intelligence (AI) systems in high-stakes domains raise concerns about proxy discrimination, unfairness, and explainability. Existing audits often fail to reveal why unfairness arises, particularly when rooted in structural bias. We propose a novel framework using formal abductive explanations to explain proxy discrimination in individual AI decisions. Leveraging background knowledge, our method identifies which features act as unjustified proxies for protected attributes, revealing hidden structural biases. Central to our approach is the concept of aptitude, a task-relevant property independent of group membership, with a mapping function aligning individuals of equivalent aptitude across groups to assess fairness substantively. As a proof of concept, we showcase the framework with examples taken from the German credit dataset, demonstrating its applicability in real-world cases.

AI公平性代理歧视可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。