即使访问平等,大模型对不同人群的回应质量仍存差异。
Equal Access, Unequal Interaction: A Counterfactual Audit of LLM Fairness
- 用反事实提示测试模型在性别、年龄、国籍下的回应差异。
- 两模型拒绝率均为零,但互动质量存在系统性不公。
- 适合关注模型公平性评估与社会影响的研究者阅读。
以往关于大语言模型公平性的研究主要聚焦于拒绝响应和安全过滤等访问层面的行为。然而,获得访问权并不等于互动质量的公平。本文通过受控的公平性审计,考察在访问已获批准后,大模型对不同人口属性(年龄、性别、国籍)用户的语气、不确定性程度和语言框架的差异。采用反事实提示设计,评估GPT-4和LLaMA-3.1-70B在职业建议任务中的表现。通过拒绝分析衡量访问公平性,使用自动语言指标(情感、礼貌度、模糊表达)衡量互动质量,并采用配对统计检验评估身份相关差异。结果显示,两个模型在所有身份组别下拒绝率均为零,表明访问完全平等;但互动质量存在系统性差异:GPT-4对年轻男性用户表现出显著更高的模糊表达,而LLaMA则在各身份组间呈现更广泛的情感波动。这表明,即便访问平等,互动层面的不公平依然存在,亟需超越仅基于拒绝行为的评估范式。
原文摘要 · Abstract (English)
Prior work on fairness in large language models (LLMs) has primarily focused on access-level behaviors such as refusals and safety filtering. However, equitable access does not ensure equitable interaction quality once a response is provided. In this paper, we conduct a controlled fairness audit examining how LLMs differ in tone, uncertainty, and linguistic framing across demographic identities after access is granted. Using a counterfactual prompt design, we evaluate GPT-4 and LLaMA-3.1-70B on career advice tasks while varying identity attributes along age, gender, and nationality. We assess access fairness through refusal analysis and measure interaction quality using automated linguistic metrics, including sentiment, politeness, and hedging. Identity-conditioned differences are evaluated using paired statistical tests. Both models exhibit zero refusal rates across all identities, indicating uniform access. Nevertheless, we observe systematic, model-specific disparities in interaction quality: GPT-4 expresses significantly higher hedging toward younger male users, while LLaMA exhibits broader sentiment variation across identity groups. These results show that fairness disparities can persist at the interaction level even when access is equal, motivating evaluation beyond refusal-based audits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。