arXiv:2606.10475cs.MAcs.AI2026-06中稿 · publication in the…

让智能体辩论更稳定:通过分离思考与表达,抗干扰能力提升至95%以上。

Decoupling Thought from Speech: Knowledge-Grounded Counterfactual Reasoning for Resilient Multi-Agent Argumentation

  • 分两阶段设计:私有规划层与公开执行层解耦,防止逻辑崩溃
  • 在270次扰动实验中,95%以上未出现严重质量下降,平均得分从0.694升至0.822
  • 适合研究多智能体系统稳定性或对抗性推理的学者

多智能体辩论框架虽能提升大模型在收敛任务中的表现,但当前优化侧重最终输出准确率,忽视过程稳定性。长期交互中,受持续扰动影响,系统常出现逻辑退化、论点重复和角色偏离。为结构化防止身份丢失并维持过程一致性,本文提出知识引导的反事实推理(KG-CFR),采用双阶段架构,严格分离私有检索增强规划缓冲区与公开执行层。我们在动态不确定性资源分配(DRAU)环境中评估该系统,引入多样性以区别于标准辩论设置。在270个完全因子化的危机模拟轨迹中,包含随机环境冲击,KG-CFR使超过95%的扰动运行避免了裁判检测到的关键后冲击质量退化(定义为质量变化Δ≤-0.20),整体论点质量从0.694提升至0.822。主要贡献在于证明架构解耦是系统在持续压力下提升韧性且不损失质量的重要因素。此外,我们引入定制向量度量,用于话语分歧与计划-执行对齐,提供方向一致的操作稳定性证据。消融实验表明,恰当的教义锚定对论点质量的影响可与前瞻性规划同等重要。初步评估显示,KG-CFR有效降低语义循环,保持智能体与原始计划的一致性。

原文摘要 · Abstract (English)

Multi-agent debate frameworks have been shown to improve large language model performance in convergent tasks, but they are currently optimized in a way that heavily favors final output accuracy rather than stability of the process. During long-horizon exchanges reactive systems under sustained perturbations often experience logic degradation, argument repetition, and role drift. To structurally prevent the identity loss and maintain the process fidelity, we introduce Knowledge-Grounded Counterfactual Reasoning (KG-CFR), a dual-stage architecture that enforces a strict separation of concerns between a private, retrieval-augmented planning buffer, and a public execution layer. We assess this system in Dynamic Resource Allocation under Uncertainty (DRAU), a dedicated 1v1v1 environment, introducing diversity as distinct from standard debate settings. Over 270 completely factorial crisis simulation trajectories with stochastic environmental shocks, KG-CFR prevents judge-detected critical post-shock degradation (defined as a quality shift, $Δ\le -0.20$) in more than 95% of perturbed runs, increasing the overall argument quality from 0.694 to 0.822. Our primary contribution is the demonstration of architectural decoupling being an important factor of systemic resilience enhancement under sustained pressure without quality loss. Furthermore, we introduce custom vector metrics for discourse divergence and plan-execution alignment that provide strong, directionally consistent evidence of operational stability. Our ablation experiments suggest that the proper doctrinal grounding can be an equally important factor for argument quality, as the prospective planning. KG-CFR, according to our initial metric evaluations, reduces semantic looping, by preserving the agent's consistency with the original plan.

多智能体辩论系统稳定性反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。