提出生成式AI中因果偏见检测新方法,可分解不同因果路径的影响。
Causal Bias Detection in Generative Artificial Intelligence
- 基于因果推断框架,统一生成式AI与传统机器学习的公平性分析
- 首次实现对生成模型替代真实因果机制的偏见量化,覆盖多路径影响
- 适用于大语言模型的种族与性别偏见分析,适合关注AI伦理的研究者
基于人工智能的自动化系统在高风险领域日益普及,引发对公平性及现实社会中人口差异持续存在的担忧。因果推断为公平性问题提供了严谨的理论框架,能将观察到的差异与潜在机制关联,并契合人类直觉和法律对歧视的定义。以往研究主要集中在标准机器学习场景,即决策者构建单一预测函数 $f_{\ ext{hat}Y}$ 预测结果变量 $Y$,而其他协变量的因果机制来自真实世界。然而,生成式AI环境更为复杂:生成模型可从任意变量集合的条件分布采样,隐式构建自身对所有因果机制的信念,而非仅学习单一预测函数。这一根本差异要求因果公平性方法的革新。本文形式化了生成式AI中的因果公平性问题,将其与标准机器学习设置统一于同一理论框架下。进一步推导出新的因果分解结果,实现对公平性影响在(a)不同因果路径上以及(b)真实机制被生成模型机制替代时的精细化量化。建立了识别条件并提出高效估计器,用于关键因果量的计算,并通过分析大语言模型在多个数据集中的种族与性别偏见,验证了该方法的有效性。
原文摘要 · Abstract (English)
Automated systems built on artificial intelligence (AI) are increasingly deployed across high-stakes domains, raising critical concerns about fairness and the perpetuation of demographic disparities that exist in the world. In this context, causal inference provides a principled framework for reasoning about fairness, as it links observed disparities to underlying mechanisms and aligns naturally with human intuition and legal notions of discrimination. Prior work on causal fairness primarily focuses on the standard machine learning setting, where a decision-maker constructs a single predictive mechanism $f_{\widehat Y}$ for an outcome variable $Y$, while inheriting the causal mechanisms of all other covariates from the real world. The generative AI setting, however, is markedly more complex: generative models can sample from arbitrary conditionals over any set of variables, implicitly constructing their own beliefs about all causal mechanisms rather than learning a single predictive function. This fundamental difference requires new developments in causal fairness methodology. We formalize the problem of causal fairness in generative AI and unify it with the standard ML setting under a common theoretical framework. We then derive new causal decomposition results that enable granular quantification of fairness impacts along both (a) different causal pathways and (b) the replacement of real-world mechanisms by the generative model's mechanisms. We establish identification conditions and introduce efficient estimators for causal quantities of interest, and demonstrate the value of our methodology by analyzing race and gender bias in large language models across different datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。