让AI模型像大脑一样高效辩论,解决冗长对话和自大主导问题。
CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate
- 构建稀疏辩论图,只让相关模型互动,减少信息过载。
- 引入信任评估模块,优化辩论路径,提升决策质量。
- 适用于需要多智能体协作的复杂推理任务,如科学问答。
当前单一大型语言模型(LLM)面临幻觉和推理能力不足的问题。为缓解此问题,多智能体辩论(MAD)作为一种有效策略应运而生,使多个LLM智能体在任务上进行深入辩论。然而,现有MAD方法存在两大问题:(a) 输入上下文过长,导致模型迷失在大量信息中,性能下降;(b) 过度自信困境,自我确信的智能体主导辩论,降低辩论有效性。为此,我们提出一种新型MAD方法——CortexDebate。受人类大脑皮层区域间通过白质形成稀疏且动态优化网络的启发,CortexDebate构建了智能体间的稀疏辩论图,每个智能体仅与对其有益的智能体辩论。为优化该图,我们提出名为麦肯锡辩论物质(MDM)的模块,作为白质的人工类比。通过集成社会学中公认的可信度衡量标准——麦肯锡信任公式,MDM实现可信评估,指导图结构优化。在四个任务类型、八个数据集上的大量实验验证了CortexDebate的有效性。
原文摘要 · Abstract (English)
Nowadays, single Large Language Model (LLM) struggles with critical issues such as hallucination and inadequate reasoning abilities. To mitigate these issues, Multi-Agent Debate (MAD) has emerged as an effective strategy, where LLM agents engage in in-depth debates with others on tasks. However, existing MAD methods face two major issues: (a) too lengthy input contexts, which causes LLM agents to get lost in plenty of input information and experiences performance drop; and (b) the overconfidence dilemma, where self-assured LLM agents dominate the debate, leading to low debating effectiveness. To address these limitations, we propose a novel MAD method called "CortexDebate". Inspired by the human brain's tendency to establish a sparse and dynamically optimized network among cortical areas governed by white matter, CortexDebate constructs a sparse debating graph among LLM agents, where each LLM agent only debates with the ones that are helpful to it. To optimize the graph, we propose a module named McKinsey-based Debate Matter (MDM), which acts as an artificial analog to white matter. By integrating the McKinsey Trust Formula, a well-established measure of trustworthiness from sociology, MDM enables credible evaluations that guide graph optimization. The effectiveness of our CortexDebate has been well demonstrated by extensive experimental results across eight datasets from four task types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。