arXiv:2608.25937cs.AIcs.MA2026-08

研究大模型在多智能体系统中如何选择答案,发现选对关键在于判断力与频率结合。

Candidate supply and answer selection shape the value of LLM judging in multi-agent systems

  • 将多智能体推理视为生成-通信-选择的进化流程,揭示无质量控制易导致错误传播。
  • 63.82%准确率提升至70.82%-70.95%,通过结合答案频率与判别信号实现。
  • 适合关注多智能体系统可靠性、模型评判机制的研究者阅读。

多智能体系统(MAS)有时已有正确答案,却仍输出错误结果。解释这一现象困难,因生成、通信与最终选择规则常同时变化。本文将多智能体推理视为候选生成、同伴沟通与终端选择的演化流水线,其中缺乏质量控制的共识可能引发类膜因漂移现象。研究两个问题:(1)大模型裁判何时能有效提供答案正确性信号;(2)该信号能否提升报告答案的准确性。我们分析了来自MMLU-Pro、GPQA、MedXpertQA和MuSR的15,336个问题,人类最后考试单独分析。为验证规则,我们在五个基准上重播了81,390个固定候选池,覆盖16,278个问题。结果有三:(1)正确答案常存在于候选集中,但系统仍会收敛于错误答案;(2)裁判可靠性非模型固有属性,而是随任务、生成器及正确答案稀有度变化;(3)结合答案频率与裁判评估仅改变最终选择规则,便使准确率从63.82%提升至70.82%-70.95%,主要挽救了被错误答案淹没的正确答案。本研究表明,在所考察系统中,生成更多候选的价值取决于其是否让正确答案出现、频繁或可识别。通过分离生成、识别与选择环节,这些发现为设计能保护正确答案的多智能体架构提供了诊断基础。

原文摘要 · Abstract (English)

Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is difficult because generation, communication and final answer-selection rules usually change simultaneously. We conceptualize multi-agent reasoning as an evolutionary pipeline of candidate generation, peer communication and terminal selection, wherein consensus without quality control can exhibit patterns of memetic drift. We study two questions: (1) when an LLM judge provides effective selection pressure by supplying a signal of answer correctness for candidates generated in a multi-agent system, and (2) when using that signal improves the reported answer. To map judge reliability, we analysed 15,336 questions from MMLU-Pro, GPQA, MedXpertQA and MuSR, with Humanity's Last Exam analysed separately. To test these rules, we replayed 81,390 fixed candidate pools drawn from 16,278 questions across five benchmarks. We report three findings. (1) A correct answer is often already present among the generated candidates, but the system can still converge on and report a wrong answer. (2) Judge reliability is not a fixed trait of the model, but varies with the task, the generator and how rare the correct answer is. (3) Combining answer frequency with the judge's evaluation changed only the final answer-selection rule and raised accuracy from 63.82% to 70.82-70.95%, primarily by rescuing correct answers that were outnumbered by popular errors. In the systems studied here, the value of generating more candidates depends on whether those extra samples make correct answers present, frequent or recognisable. By isolating generation, recognition and selection, these findings establish a diagnostic basis for designing multi-agent architectures that protect generated correct answers from being lost.

多智能体大模型判别答案选择系统优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。