揭露多智能体系统中大模型隐性偏见,提出可量化评估新基准
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
- 构建场景化评测框架,通过多智能体打分检测偏见
- 发现偏见缓解可能更倾向弱势群体而非绝对中立
- 适合关注公平性与评估基准的研究者使用
多智能体系统由多个在共享环境中互动的AI模型构成,广泛用于基于角色的交互。然而,若设计不当,这些系统可能强化大语言模型(LLMs)中的隐性偏见,引发公平性与代表性问题。我们提出MALIBU,一个新型基准,用于评估基于LLM的多智能体系统在多大程度上隐性强化社会偏见与刻板印象。MALIBU通过情景化评估方法,让AI模型在预设情境中完成任务,其响应由基于LLM的多智能体评判系统分两阶段评估:第一阶段,评判者对标注特定人口属性(如性别、种族、宗教)的响应在四个维度上打分;第二阶段,评判者对比分配不同人口属性的成对响应,评分并选择更优者。研究量化了LLM生成内容中的偏见,发现偏见缓解策略可能更倾向于弱势群体而非绝对中立,强调需采用细致的检测方法、平衡的公平策略以及透明的评估基准来优化多智能体系统。
原文摘要 · Abstract (English)
Multi-agent systems, which consist of multiple AI models interacting within a shared environment, are increasingly used for persona-based interactions. However, if not carefully designed, these systems can reinforce implicit biases in large language models (LLMs), raising concerns about fairness and equitable representation. We present MALIBU, a novel benchmark developed to assess the degree to which LLM-based multi-agent systems implicitly reinforce social biases and stereotypes. MALIBU evaluates bias in LLM-based multi-agent systems through scenario-based assessments. AI models complete tasks within predefined contexts, and their responses undergo evaluation by an LLM-based multi-agent judging system in two phases. In the first phase, judges score responses labeled with specific demographic personas (e.g., gender, race, religion) across four metrics. In the second phase, judges compare paired responses assigned to different personas, scoring them and selecting the superior response. Our study quantifies biases in LLM-generated outputs, revealing that bias mitigation may favor marginalized personas over true neutrality, emphasizing the need for nuanced detection, balanced fairness strategies, and transparent evaluation benchmarks in multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。