arXiv:2410.22932cs.CL2024-10被引 33

多智能体对话系统在复杂任务中表现更优,但简单任务易出错。

Multi-Agent Large Language Models for Conversational Task-Solving

  • 用多个角色智能体模拟讨论,通过分工协作提升复杂推理能力。
  • 复杂任务下多智能体优于单模型,但简单任务中错误率更高。
  • 适合研究对话效率、智能体协作机制与系统安全的学者参考。

在单个大语言模型主导人工智能领域多年后,多智能体系统成为对话任务求解的新主角。尽管已有研究展示其在推理和创意任务中的潜力,但对其在对话范式下的局限性及个体智能体影响的分析仍缺失。本文系统评估了多智能体系统在不同讨论范式下的表现,涵盖生成任务与问答任务。通过分析2022至2024年的20项研究,提出一个分类体系,并引入一套部署框架。实验表明,多智能体在复杂推理任务中表现优异,通过专家角色分工超越单模型;但在基础任务中表现不佳。具体发现三类问题:1)更长的讨论虽提升推理能力,但导致任务偏离,问题漂移,因此短对话更适合基础任务;2)长时间讨论可能引发对齐崩溃,带来新的安全风险;3)长生成过程出现讨论垄断,影响摘要等任务中的决策公平性。本研究揭示了多智能体交互在不同对话模式下的潜力与挑战,为未来提升效率、性能与安全性提供洞见。

原文摘要 · Abstract (English)

In an era where single large language models have dominated the landscape of artificial intelligence for years, multi-agent systems arise as new protagonists in conversational task-solving. While previous studies have showcased their potential in reasoning tasks and creative endeavors, an analysis of their limitations concerning the conversational paradigms and the impact of individual agents is missing. It remains unascertained how multi-agent discussions perform across tasks of varying complexity and how the structure of these conversations influences the process. To fill that gap, this work systematically evaluates multi-agent systems across various discussion paradigms, assessing their strengths and weaknesses in both generative tasks and question-answering tasks. Alongside the experiments, I propose a taxonomy of 20 multi-agent research studies from 2022 to 2024, followed by the introduction of a framework for deploying multi-agent LLMs in conversational task-solving. I demonstrate that while multi-agent systems excel in complex reasoning tasks, outperforming a single model by leveraging expert personas, they fail on basic tasks. Concretely, I identify three challenges that arise: 1) While longer discussions enhance reasoning, agents fail to maintain conformity to strict task requirements, which leads to problem drift, making shorter conversations more effective for basic tasks. 2) Prolonged discussions risk alignment collapse, raising new safety concerns for these systems. 3) I showcase discussion monopolization through long generations, posing the problem of fairness in decision-making for tasks like summarization. This work uncovers both the potential and challenges that arise with multi-agent interaction and varying conversational paradigms, providing insights into how future research could improve the efficiency, performance, and safety of multi-agent LLMs.

多智能体对话系统大模型推理任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。