MALLM框架可系统配置多智能体辩论,提升集体决策能力。
MALLM: Multi-Agent Large Language Models Framework
- 通过144种组合配置智能体角色、生成方式、讨论模式和决策规则
- 支持加载Hugging Face数据集并自动评估不同配置效果
- 适合研究多智能体协作机制的学者与开发者使用
多智能体辩论(MAD)通过扩展推理计算和利用专业能力,展现出增强集体智能的潜力。现有MAD框架多聚焦工具使用,缺乏集成评估能力,且对智能体人格、响应生成器、讨论范式和决策协议的可配置性有限。我们提出MALLM(多智能体大语言模型框架),一个开源系统,支持对MAD组件的系统化分析。MALLM提供超过144种独特的MAD配置,涵盖智能体人格(如专家、人格化)、响应生成器(如批判性、推理型)、讨论范式(如记忆型、接力型)和决策协议(如投票、共识)。MALLM通过简单配置文件定义辩论流程,并支持加载任意Hugging Face文本数据集(如MMLU-Pro、WinoGrande),内置评估流水线,便于对比不同配置的表现。该框架使研究人员能系统地配置、运行和评估辩论,推动对各组件及其相互作用的理解。
原文摘要 · Abstract (English)
Multi-agent debate (MAD) has demonstrated the ability to augment collective intelligence by scaling test-time compute and leveraging expertise. Current frameworks for multi-agent debate are often designed towards tool use, lack integrated evaluation, or provide limited configurability of agent personas, response generators, discussion paradigms, and decision protocols. We introduce MALLM (Multi-Agent Large Language Models), an open-source framework that enables systematic analysis of MAD components. MALLM offers more than 144 unique configurations of MAD, including (1) agent personas (e.g., Expert, Personality), (2) response generators (e.g., Critical, Reasoning), (3) discussion paradigms (e.g., Memory, Relay), and (4) decision protocols (e.g., Voting, Consensus). MALLM uses simple configuration files to define a debate. Furthermore, MALLM can load any textual Hugging Face dataset (e.g., MMLU-Pro, WinoGrande) and provides an evaluation pipeline for easy comparison of MAD configurations. MALLM enables researchers to systematically configure, run, and evaluate debates for their problems, facilitating the understanding of the components and their interplay.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。