arXiv:2410.04452cs.CLcs.AI2024-10中稿 · presentation at th…被引 5

用多智能体系统检测大模型认知偏见,效果比GPT-4提升35.1%。

MindScope: Exploring cognitive biases in large language models through Multi-Agent Systems

  • 构建包含5170题的静态数据集与可动态交互的多智能体框架。
  • 提出融合RAG与对抗辩论的检测方法,准确率提升35.10%。
  • 适合研究模型偏见、心理实验或可解释AI的科研人员使用。

检测大语言模型中的认知偏见是一项重要任务,旨在揭示模型内部存在的认知偏差。现有方法普遍存在检测能力不全、可识别偏见类型有限的问题。为此,我们提出了'MindScope'数据集,其独特地融合了静态与动态元素:静态部分包含5170个开放式问题,覆盖72类认知偏见;动态部分采用基于规则的多智能体通信框架,支持多轮对话生成,灵活适用于多种涉及大模型的心理学实验。此外,我们设计了一种多智能体检测方法,整合检索增强生成(RAG)、竞争性辩论及强化学习决策模块,广泛适用于各类检测任务。实验表明,该方法相较GPT-4检测准确率最高提升35.10%。代码与附录详见https://github.com/2279072142/MindScope。

原文摘要 · Abstract (English)

Detecting cognitive biases in large language models (LLMs) is a fascinating task that aims to probe the existing cognitive biases within these models. Current methods for detecting cognitive biases in language models generally suffer from incomplete detection capabilities and a restricted range of detectable bias types. To address this issue, we introduced the 'MindScope' dataset, which distinctively integrates static and dynamic elements. The static component comprises 5,170 open-ended questions spanning 72 cognitive bias categories. The dynamic component leverages a rule-based, multi-agent communication framework to facilitate the generation of multi-round dialogues. This framework is flexible and readily adaptable for various psychological experiments involving LLMs. In addition, we introduce a multi-agent detection method applicable to a wide range of detection tasks, which integrates Retrieval-Augmented Generation (RAG), competitive debate, and a reinforcement learning-based decision module. Demonstrating substantial effectiveness, this method has shown to improve detection accuracy by as much as 35.10% compared to GPT-4. Codes and appendix are available at https://github.com/2279072142/MindScope.

认知偏见多智能体检测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。