让多个AI助手的思考过程互相补全,比只看一致答案更准。
Beyond Consensus: Trace-Level Synthesis in Mixture of Agents

- 用完整推理链替代答案投票,实现细粒度融合
- 在数学与科学任务中,正确率超越多模型组合
- 通过扰动生成多样性路径,避免盲目依赖共识
当多个大语言模型代理解决同一问题时,传统做法将各自推理压缩为多数投票或分层合成,以意见一致为终点。我们发现此举过度损失信息:一个读取完整推理链的LLM聚合器,即使在所有代理一致时也能恢复正确解,且有益修正始终超过有害偏差——即“聚合悖论”。多数投票存在天花板,扰动多样性无法突破(错误相关性相同);聚合优势源于推理链层面的互补性,能从少数路径中拼出正确中间步骤。这催生了自洽混合代理(Self-Consistent Mixture of Agents),通过语义保持的输入扰动生成推理多样性,借助锚定精炼保障多数结果不退化,并始终合成,永不基于共识截断。单一模型经扰动后表现优于异构模型池,在结构化推理、博士级科学、竞赛数学与编程竞赛中均胜出。聚合单位应是推理链,而非答案。
原文摘要 · Abstract (English)
When multiple LLM agents solve the same problem, standard practice compresses each agent's reasoning into a majority vote or layered synthesis, treating agreement as the finish line. We show this is unnecessarily lossy: an LLM aggregator that reads complete reasoning traces recovers correct solutions even when agents unanimously agree, with beneficial corrections consistently outweighing harmful ones -- the \emph{aggregation paradox}. Majority voting has a ceiling that perturbation diversity does not raise (error correlations are identical); the aggregator's gain comes from trace-level complementarity, assembling correct intermediate steps from minority chains that voting discards. These findings motivate Self-Consistent Mixture of Agents which generates trace diversity through semantic-preserving input perturbations, safeguards the majority via anchored refinement with provable non-degradation guarantees, and always synthesizes -- never gates on consensus. A single model with perturbation-induced trace variation outperforms heterogeneous model pools across structured reasoning, PhD-level science, competition mathematics, and competitive programming. The unit of aggregation should be the reasoning trace, not the answer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。