研究发现,角色设定会引发多智能体系统中的偏见,影响信任与坚持度。
From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent Interactions
- 通过控制实验分析角色设定对智能体行为的影响。
- 优势群体角色被低估信任度且更少坚持己见。
- 智能体存在同群偏好,倾向服从相同角色者。
基于大语言模型的多智能体系统被广泛用于模拟人类互动和协作任务。常见做法是为智能体分配角色以促进行为多样性,但这一做法可能引入偏见。本文系统研究了角色设定在多智能体互动中引发的偏见,重点关注信任度(他人接受其观点的程度)和坚持度(坚持表达观点的强度)。通过一系列协作解决问题和说服任务的控制实验,发现:(1) 大语言模型智能体在信任度和坚持度上均存在偏见,历史上占优势群体的角色(如男性、白人)反而被认为更不可信且更少坚持己见;(2) 智能体表现出显著的内群体偏好,更倾向于服从具有相同角色的其他智能体。这些偏见在不同大模型、群体规模及交互轮次下持续存在,凸显了提升多智能体系统公平性与可靠性的重要性和紧迫性。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based multi-agent systems are increasingly used to simulate human interactions and solve collaborative tasks. A common practice is to assign agents with personas to encourage behavioral diversity. However, this raises a critical yet underexplored question: do personas introduce biases into multi-agent interactions? This paper presents a systematic investigation into persona-induced biases in multi-agent interactions, with a focus on social traits like trustworthiness (how an agent's opinion is received by others) and insistence (how strongly an agent advocates for its opinion). Through a series of controlled experiments in collaborative problem-solving and persuasion tasks, we reveal that (1) LLM-based agents exhibit biases in both trustworthiness and insistence, with personas from historically advantaged groups (e.g., men and White individuals) perceived as less trustworthy and demonstrating less insistence; and (2) agents exhibit significant in-group favoritism, showing a higher tendency to conform to others who share the same persona. These biases persist across various LLMs, group sizes, and numbers of interaction rounds, highlighting an urgent need for awareness and mitigation to ensure the fairness and reliability of multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。