通过角色工程揭示大模型在不同社会身份下的隐性偏见动态变化
Invisible Influences: Investigating Implicit Intersectional Biases through Persona Engineering in Large Language Models

- 设计新指标BADx,结合动态角色上下文评估偏见放大与可解释性
- 五款主流模型在不同角色下表现差异显著,部分模型偏见随角色剧烈波动
- 适用于关注模型公平性、偏见检测的研究者与AI伦理实践者
大型语言模型虽能生成类人语言,却常内嵌并放大隐性交叉偏见,尤其在角色驱动情境中。现有审计方法依赖静态嵌入测试(CEAT、I-WEAT、I-SEAT),仅量化绝对关联强度,难以捕捉角色变化带来的动态偏见。本文提出全新可扩展指标BADx,用于测量角色诱导的偏见放大,并融合局部可解释性分析。BADx包含三部分:基于经典测试的差分偏见得分(BAD)、角色敏感度指数(PSI)与波动性(标准差),辅以LIME分析增强可解释性。研究分为两部分:任务1建立静态偏见基线;任务2在六种角色框架(边缘化与结构优势型)下评估五款先进模型(GPT-4o、DeepSeek-R1、LLaMA-4、Claude 4.0 Sonnet、Gemma-3n E4B)的BADx、PSI与波动性。结果表明,角色上下文显著影响偏见表现:GPT-4o敏感度高且波动剧烈;DeepSeek-R1抑制偏见但波动不稳;LLaMA-4波动低、偏见稳定且放大有限;Claude 4.0 Sonnet调节平衡;Gemma-3n E4B波动最低,放大中等。相比静态方法,BADx更有效揭示被忽略的上下文敏感偏见。
原文摘要 · Abstract (English)
Large Language Models (LLMs) excel at human-like language generation but often embed and amplify implicit, intersectional biases, especially under persona-driven contexts. Existing bias audits rely on static, embedding-based tests (CEAT, I-WEAT, I-SEAT) that quantify absolute association strengths. We show that they have limitations in capturing dynamic shifts when models adopt social roles. We address this gap by introducing the Bias Amplification Differential and Explainability Score (BADx): a novel, scalable metric that measures persona-induced bias amplification and integrates local explainability insights. BADx comprises three components - differential bias scores (BAD, based on CEAT, I-WEAT, I-SEAT),Persona Sensitivity Index (PSI), and Volatility (Standard Deviation), augmented by LIME-based analysis for emphasizing explainability. This study is divided and performed as two different tasks. Task 1 establishes static bias baselines, and Task 2 applies six persona frames (marginalized and structurally advantaged) to measure BADx, PSI, and volatility. This is studied across five state-of-the-art LLMs (GPT-4o, DeepSeek-R1, LLaMA-4, Claude 4.0 Sonnet and Gemma-3n E4B). Results show persona context significantly modulates bias. GPT-4o exhibits high sensitivity and volatility; DeepSeek-R1 suppresses bias but with erratic volatility; LLaMA-4 maintains low volatility and a stable bias profile with limited amplification; Claude 4.0 Sonnet achieves balanced modulation; and Gemma-3n E4B attains the lowest volatility with moderate amplification. BADx performs better than static methods by revealing context-sensitive biases overlooked in static methods. Our unified method offers a systematic way to detect dynamic implicit intersectional bias in five popular LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。