研究大模型解释决策时对不同人群的公平性差异,发现解释质量不均。
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups

- 构建五维解释公平性框架,量化解释在深度、语气、用词等方面差异
- 80个提示模板下,所有模型在5个维度均存在显著不公平,部分差距超5倍
- 模型预训练数据影响解释风格,仅靠提示无法消除语言风格不公
大型语言模型(LLMs)不仅用于做决策,还被要求提供解释。尽管决策公平性受广泛研究,但模型对不同人口群体是否以同等质量、深度、语气和语言复杂度进行解释,仍缺乏关注。本文提出解释公平性分类体系(EFT),包含五个可操作的维度:冗长度差异、情感倾向差异、认知模糊性差异、与决策关联的解释差异、词汇复杂度差异。在80个提示模板、4个关键决策领域(招聘、医疗分诊、信贷评估、法律裁决)及5个模型(GPT-4.1、Claude Sonnet、LLaMA 3.3 70B、GPT-OSS 120B、Qwen3 32B)中开展实证研究。引入两个新黑盒指标:模糊密度得分(HDS)与解释忠实性代理(EFP)。在最多400组提示对中,所有8个EFT指标均显示统计显著差异(Cohen's d从微小到巨大,所有p_BH < 10^(-62))。模型选择强烈影响差异幅度:Qwen3 32B的冗长度差异是LLaMA 3.3 70B的5.9倍。两种提示缓解策略显著降低EFP差异(78%-95%),但对风格维度无效,表明风格不公源自预训练分布,非仅通过部署指令可解决。研究提供可复现的解释级公平性审计框架,对人工智能监管与部署具有重要启示。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed not only to make decisions but to explain them. While AI decision fairness has been studied extensively, the fairness of AI explanations (whether LLMs justify decisions with equal quality, depth, tone, and linguistic sophistication across demographic groups) has received little attention. This paper introduces the Explanation Fairness Taxonomy (EFT), a framework comprising five formally defined, operationalizable dimensions: Verbosity Disparity, Sentiment Disparity, Epistemic Hedging Disparity, Decision-Linked Explanation Disparity, and Lexical Complexity Disparity. The taxonomy is instantiated in a controlled empirical study across 80 prompt templates, four consequential decision domains (hiring, medical triage, credit assessment, legal judgment), and five LLMs: GPT-4.1, Claude Sonnet, LLaMA 3.3 70B, GPT-OSS 120B, and Qwen3 32B. Two novel black-box metrics are introduced: the Hedging Density Score (HDS) and the Explanation Faithfulness Proxy (EFP), a heuristic indicator of decision-linked explanation variation. Across up to 400 prompt pairs, all eight EFT metrics show statistically significant disparities (Cohen's d ranging from small to large, all p_BH < 10^(-62)). Model choice is strongly associated with disparity magnitude: Qwen3 32B exhibits verbosity disparities 5.9x larger than LLaMA 3.3 70B. Two prompting-based mitigations show significant reductions in EFP disparity (78-95%) but no significant effect on stylistic dimensions, consistent with the hypothesis that stylistic explanation inequalities are encoded in pre-training distributions and are not resolvable through deployment-level instruction alone. A reproducible measurement framework is offered for explanation-level fairness auditing, with implications for AI regulation and deployment practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。