用多智能体辩论消除大模型文化偏见,效果显著提升。
Mitigating Cultural Bias in LLMs via Multi-Agent Cultural Debate
- 设计多智能体文化辩论框架,赋予智能体明确文化身份进行讨论。
- 在中英双语评测集上实现57.6%无偏率(LLM判别)和86.0%(自建判别器)。
- 无需训练即可跨文化泛化,适合关注AI公平性的研究者使用。
大语言模型存在系统性西方中心偏见,但以非西方语言(如中文)提示能否缓解该问题仍缺乏研究。现有评估方法强制输出归入预设文化类别,缺乏中立选项;而缓解策略依赖昂贵的多文化语料或功能型代理框架(如Planner--Critique),未显式体现文化身份。为此,我们提出CEBiasBench中英双语评测基准与Multi-Agent Vote(MAV)机制,支持显式的“无偏”判断。实验发现,中文提示仅将偏见转向东亚视角,并未消除。为解决此问题,我们提出训练免费的多智能体文化辩论(MACD)框架:为代理分配不同文化身份,采用“求同存异”策略引导协商。在CEBiasBench上,MACD以GPT-4o为基线时,获57.6%平均无偏率(由LLM判别)与86.0%(由MAV判别),优于基线(47.6%与69.0%),且在阿拉伯语CAMeL基准上表现良好,证明显式文化表征对跨文化公平至关重要。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit systematic Western-centric bias, yet whether prompting in non-Western languages (e.g., Chinese) can mitigate this remains understudied. Answering this question requires rigorous evaluation and effective mitigation, but existing approaches fall short on both fronts: evaluation methods force outputs into predefined cultural categories without a neutral option, while mitigation relies on expensive multi-cultural corpora or agent frameworks that use functional roles (e.g., Planner--Critique) lacking explicit cultural representation. To address these gaps, we introduce CEBiasBench, a Chinese--English bilingual benchmark, and Multi-Agent Vote (MAV), which enables explicit ``no bias'' judgments. Using this framework, we find that Chinese prompting merely shifts bias toward East Asian perspectives rather than eliminating it. To mitigate such persistent bias, we propose Multi-Agent Cultural Debate (MACD), a training-free framework that assigns agents distinct cultural personas and orchestrates deliberation via a "Seeking Common Ground while Reserving Differences" strategy. Experiments demonstrate that MACD achieves 57.6% average No Bias Rate evaluated by LLM-as-judge and 86.0% evaluated by MAV (vs. 47.6% and 69.0% baseline using GPT-4o as backbone) on CEBiasBench and generalizes to the Arabic CAMeL benchmark, confirming that explicit cultural representation in agent frameworks is essential for cross-cultural fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。