通过分析智能体弱点,动态干预提升复杂问题求解可靠性。
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
- 用控制理论方法为执行智能体建立性能指纹,识别其固有缺陷。
- 在GAIA数据集上,系统准确率显著高于单智能体和普通协作系统。
- 适合关注智能体协同与可信推理的研究者与开发者参考。
大型语言模型的发展使智能体能够调用外部工具解决复杂现实问题,但长上下文和噪声工具输出会降低系统可靠性。为此,我们提出AWorld框架下的动态多智能体系统(MAS),由监督型守护智能体对执行智能体进行动态干预,验证并修正推理过程,从而提升鲁棒性。为进一步超越通用监督,我们借鉴控制理论中的系统辨识方法:先在基准数据集上离线分析执行智能体的性能表现,构建其独有的「性能指纹」;在线运行时,守护智能体利用该指纹实施针对性干预,基于已知失败模式而非仅响应即时逻辑错误。在GAIA数据集上的大量实验表明,该方法显著提升了系统有效性和稳定性,不仅优于单智能体系统,也优于未经优化的多智能体方案。该系统在权威GAIA排行榜中位列开源项目第一。研究结果表明,构建真正可信的智能系统,不仅需要协作,更需基于实证对每个智能体的能力与局限有深入理解。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has empowered intelligent agents to leverage diverse external tools for solving complex real-world problems. However, this reliance introduces new challenges, as extended contexts and noisy tool outputs can undermine system reliability. To address this, we propose a dynamic Multi-Agent System (MAS) in our AWorld framework, where an Execution Agent is supervised by a Guard Agent that provides on-demand dynamic maneuvering, verifying and correcting the reasoning process to improve robustness over single-agent systems. To move beyond this generic supervision, we enhance the architecture with a methodology inspired by System Identification from control theory. This method first profiles the Execution Agent offline on a benchmark dataset to create a "performance fingerprint" of its unique weaknesses. The Guard Agent then leverages this fingerprint online to deliver profile-aware supervision, making targeted interventions based on known failure patterns rather than merely reacting to immediate logical flaws. Extensive experiments on the GAIA dataset demonstrate that this profile-aware MAS significantly improves both effectiveness and stability, outperforming not only single-agent systems but also its naive counterpart. This superior performance led our system to achieve first place among open-source projects on the prestigious GAIA leaderboard. These findings highlight that building truly trustworthy intelligent systems requires not just collaboration, but a deep, empirically-grounded understanding of each agent's unique capabilities and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。