arXiv:2509.21080cs.CLcs.AI2025-09

发现大模型生成访谈脚本时存在文化偏见,提出新方法有效缓解。

InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation

  • 构建4000条提示的InsideOut基准,评估模型对10种文化的立场偏向
  • 5个主流模型在非西方文化中超过88%采用'外人'视角,严重失衡
  • 基于智能体的缓解框架显著降低文化偏差,效果优于传统提示法

大型语言模型在故事和访谈脚本生成等下游应用中取得进展,但近期研究揭示其内容存在文化相关公平性问题。本文识别并系统研究了模型的‘内人-外人’偏见:模型在生成时倾向于将自身置于主流文化‘内人’位置,而将非主流文化边缘化为‘外人’。为此,我们提出InsideOut基准,包含4000个生成提示,通过一个文化情境化的访谈脚本生成任务(即模型作为记者采访10种不同文化背景的当地人)来量化该偏见,并设计三个评估指标。对5个先进大模型的实证评估显示,平均而言,在美国语境下超过88%的脚本呈现‘内人’语气,而在非西方文化中则过度倾向‘外人’立场。为缓解此偏见,我们提出两种推理时干预方法:基础提示法的公平干预支柱(FIP),以及由单智能体(MFA-SA)、分层智能体(MFA-HA)和自主代理规划(MFA-Plan)组成的结构化公平代理缓解框架(MFA)。实验结果表明,基于智能体的方法在缓解内人-外人偏见方面表现卓越且稳定:例如,在文化契合度差距(CAG)指标上,MFA-SA使Llama模型的偏见降低89.70%,MFA-HA使Qwen模型的偏见降低82.54%。这些发现证明了智能体方法在生成模型偏见缓解中的有效性,是未来重要方向。

原文摘要 · Abstract (English)

Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation. However, recent research raised concerns about culture-related fairness issues in LLM-generated content. In this work, we identify and systematically investigate LLMs' insider-outsider bias, a phenomenon where models position themselves as "insiders" of mainstream cultures during generation while externalizing less dominant cultures. We propose the InsideOut benchmark with 4,000 generation prompts and three evaluation metrics to quantify this bias through a culturally situated interview script generation task, in which an LLM is positioned as a reporter interviewing local people across 10 diverse cultures. Empirical evaluation on 5 state-of-the-art LLMs reveals that while models adopt insider tones in over 88% US-contexted scripts on average, they disproportionately default to "outsider" stances for non-Western cultures. To mitigate these biases, we propose 2 inference-time methods: a baseline prompt-based Fairness Intervention Pillars (FIP) method, and a structured Mitigation via Fairness Agents (MFA) framework consisting of a Single-Agent (MFA-SA), a Hierarchical-Agent (MFA-HA), and an autonomous Agentic Planning (MFA-Plan) pipeline. Empirical results demonstrate that agent-based MFA methods achieve outstanding and robust performance in mitigating the insider-outsider bias: For instance, on the Cultural Alignment Gap (CAG) metric, MFA-SA reduces bias in Llama model by 89.70 % and MFA-HA mitigates bias in Qwen by 82.54%. These findings showcase the effectiveness of agent-based methods as a promising direction for mitigating biases in generative LLMs.

大模型偏见文化公平智能体缓解访谈生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。