用大模型特性提升多智能体系统抗故障能力
Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance
- 基于大模型的智能体通过自我反思识别异常信息流
- 在85.7%故障率下仍保持高可靠性,优于传统方法
- 适合研究可信多智能体系统或安全关键场景的开发者
保障多智能体系统(MAS)中代理架构的可靠性,并在故障发生时有效识别问题代理,是关键挑战。大型语言模型(LLMs)推动了基于大模型的智能体成为主流,显著提升了复杂问题求解与世界建模能力。然而,这一转变对系统可靠性的影响尚不明确:以大模型代理替代传统代理是否能真正增强系统的鲁棒性?本文从拜占庭容错视角,探究并量化了基于大模型的智能体的可靠性。我们发现,大模型代理在处理错误消息流时表现出更强的怀疑倾向,这一特性使其在不同拓扑结构下均优于传统代理。受初步实验启发,我们提出CP-WBFT——一种基于置信度探测的加权拜占庭容错共识机制,利用大模型的内在反思与判别能力,通过探测式加权信息传输来提升系统稳定性。大量实验表明,在极端拜占庭条件下(85.7%故障率),CP-WBFT在多种网络拓扑中均表现优异,不仅在数学推理与安全评估任务中取得显著准确率,且整体可靠性远超传统方法。
原文摘要 · Abstract (English)
Ensuring the reliability of agent architectures and effectively identifying problematic agents when failures occur are crucial challenges in multi-agent systems (MAS). Advances in large language models (LLMs) have established LLM-based agents as a major branch of MAS, enabling major breakthroughs in complex problem solving and world modeling. However, the reliability implications of this shift remain largely unexplored. i.e., whether substituting traditional agents with LLM-based agents can effectively enhance the reliability of MAS. In this work, we investigate and quantify the reliability of LLM-based agents from the perspective of Byzantine fault tolerance. We observe that LLM-based agents demonstrate stronger skepticism when processing erroneous message flows, a characteristic that enables them to outperform traditional agents across different topological structures. Motivated by the results of the pilot experiment, we design CP-WBFT, a confidence probe-based weighted Byzantine Fault Tolerant consensus mechanism to enhance the stability of MAS with different topologies. It capitalizes on the intrinsic reflective and discriminative capabilities of LLMs by employing a probe-based, weighted information flow transmission method to improve the reliability of LLM-based agents. Extensive experiments demonstrate that CP-WBFT achieves superior performance across diverse network topologies under extreme Byzantine conditions (85.7\% fault rate). Notably, our approach surpasses traditional methods by attaining remarkable accuracy on various topologies and maintaining strong reliability in both mathematical reasoning and safety assessment tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。