多智能体辩论易偏离原问题,导致性能下降。
Stay Focused: Problem Drift in Multi-Agent Debate
- 分析辩论中问题漂移现象并量化其程度。
- 发现生成类任务漂移率高达76%-89%。
- 提出检测与缓解漂移的DRIFTJudge和DRIFTPolicy方法。
多智能体辩论——多个大语言模型通过回合制交互讨论问题——在解决知识与推理任务方面展现出潜力。然而,当面对需要长推理链的复杂问题时,该方法表现受限。本文分析了多轮交互中多智能体辩论逐渐偏离初始问题的现象,定义为问题漂移,并在十项任务(包括三类生成、三类知识、三类推理及一类指令遵循任务)中量化其存在。研究发现,生成类任务因答案空间主观性,漂移率高达76%-89%,而高复杂度任务漂移率仅为7%-21%。八位人类专家分析170个出现漂移的辩论案例,发现主要问题为缺乏进展(35%)、低质量反馈(26%)及表述不清(25%)。本文提出首个基准方法DRIFTJudge(基于LLM的判别器)用于检测问题漂移,并设计DRIFTPolicy以缓解31%的漂移情况。本研究揭示了长辩论可能损害性能的根本原因,为优化多智能体系统提供方向。
原文摘要 · Abstract (English)
Multi-agent debate - multiple instances of large language models discussing problems in turn-based interaction - has shown promise for solving knowledge and reasoning tasks. However, these methods show limitations when solving complex problems that require longer reasoning chains. We analyze how multi-agent debate drifts away from the initial problem over multiple turns, thus harming task performance. We define this phenomenon as problem drift and quantify its presence across ten tasks (i.e., three generative, three knowledge, three reasoning, and one instruction-following task). We find that generative tasks drift often due to the subjectivity of the answer space (76-89%), compared to high-complexity tasks (7-21%). To identify the reasons, eight human experts analyze 170 multi-agent debates suffering from problem drift. We find the most common issues related to this drift are the lack of progress (35% of cases), low-quality feedback (26% of cases), and a lack of clarity (25% of cases). We propose DRIFTJudge, an LLM-as-a-judge method, as a first baseline to detect problem drift. We also propose DRIFTPolicy, which mitigates 31% of problem drift cases. Our study is a step toward understanding a key limitation of multi-agent debate, highlighting why longer debates can harm task performance and how problem drift could be addressed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。