让多个AI辩论更高效,自动判断何时结束并给出答案。
Meta-Moderator: Empowering Multi-Agent Debate with Meta-Cognition

- 引入可学习的元认知调节器,动态控制辩论进程。
- 在五个基准上表现优于传统决策层,且跨任务泛化能力强。
- 减少冗余讨论,更智能地整合有效观点,适合复杂推理场景。
多智能体辩论可通过激发多样假设与批判提升大模型推理能力,但常受限于薄弱的调解机制。现有流程依赖固定预算、基于共识的停止策略或未经训练的裁判,导致重复讨论和不可靠的证据聚合。本文将调解视为一种元认知过程,通过监控辩论价值、控制讨论节奏并裁定最终答案,提出可学习的Meta-Moderator框架,实现辩论的动态调控与适时终止。该框架通过结果驱动的策略优化独立于辩论者进行训练,使调解成为明确的能力而非提示工程的副产品。在五个基准测试中,Meta-Moderator性能超越广泛使用的决策层,并展现出跨任务与系统配置的迁移能力。进一步分析表明,其能更精准分配辩论资源,且在出现有价值假设后有效降低错误聚合风险。
原文摘要 · Abstract (English)
Multi-agent debate can improve large language model reasoning by eliciting diverse hypotheses and critiques, yet its performance is often constrained by weak moderation. Common pipelines rely on fixed budgets, agreement-based stopping, or untrained judges, leading to redundant deliberation and unreliable evidence aggregation. We cast moderation as a meta-cognitive process, monitoring debate utility, controlling deliberation, and adjudicating a final answer, and introduce Meta-Moderator, a learnable framework that dynamically regulates debate and decides when to finalize an answer. Meta-Moderator is trained independently of the debaters via outcome-driven policy optimization, making debate regulation an explicit capability rather than an incidental effect of prompting. Across five benchmarks, Meta-Moderator outperforms widely used decision layers and transfers across tasks and system configurations. Further analyses show that it allocates debate more selectively and reduces mis-aggregation after informative hypotheses appear.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。