发现对话型AI在群组互动中会放大立场偏见,现有检测工具却看不见。
Unmasking Conversational Bias in AI Multiagent Systems
- 构建双模型对话模拟小圈子,观察立场变化
- 保守立场在封闭对话中反而更极端,违背预期
- 现有问卷式检测法无法识别此类偏见,需新工具
检测生成模型输出中的偏见对降低其在关键场景应用的风险至关重要。然而,现有方法大多将模型孤立看待,忽视其在实际情境中的应用。尤其在涉及生成式模型的多智能体系统中,由此产生的偏见仍研究不足。为此,我们提出一种框架,用于量化对话型大语言模型(LLMs)多智能体系统中的偏见。通过模拟小型回音室,让一对初始立场一致的LLMs就争议话题展开讨论。出乎意料的是,生成消息中的立场出现显著转变,尤其在所有智能体初始均持保守观点的回音室中——这与许多LLM普遍存在的自由主义倾向相一致。关键的是,这种回音室实验中显现的偏见,当前最先进的基于问卷的检测方法无法识别。这凸显了为人工智能多智能体系统开发更先进偏见检测与缓解工具的迫切需求。实验代码已公开。
原文摘要 · Abstract (English)
Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application in critical settings. However, the majority of existing methodologies for identifying biases in generated text consider the models in isolation and neglect their contextual applications. Specifically, the biases that may arise in multi-agent systems involving generative models remain under-researched. To address this gap, we present a framework designed to quantify biases within multi-agent systems of conversational Large Language Models (LLMs). Our approach involves simulating small echo chambers, where pairs of LLMs, initialized with aligned perspectives on a polarizing topic, engage in discussions. Contrary to expectations, we observe significant shifts in the stance expressed in the generated messages, particularly within echo chambers where all agents initially express conservative viewpoints, in line with the well-documented political bias of many LLMs toward liberal positions. Crucially, the bias observed in the echo-chamber experiment remains undetected by current state-of-the-art bias detection methods that rely on questionnaires. This highlights a critical need for the development of a more sophisticated toolkit for bias detection and mitigation for AI multi-agent systems. The code to perform the experiments is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。