arXiv:2505.24354cs.CL2025-05ACL被引 8

构建可复现的语言智能体研究框架,统一开发与评估流程。

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research

  • 用图结构工作流引擎实现模块化智能体架构。
  • 复杂推理虽强,但思维链方法更高效且低开销。
  • 提供标准化评测体系,适合研究人员复现实验。

由大语言模型驱动的语言智能体在理解、推理和执行复杂任务方面展现出显著能力。然而,构建可靠的智能体面临重大挑战:工程负担重、组件缺乏标准、评估框架不足。我们提出AGORA(基于图的推理与评估智能体编排系统),通过三项核心贡献解决这些问题:(1) 模块化架构,包含基于图的工作流引擎、高效的内存管理及清晰的组件抽象;(2) 一套可复用的智能体算法,实现当前最先进的推理方法;(3) 严谨的评估框架,支持多维度系统性比较。在数学推理和多模态任务上的大量实验表明,尽管复杂推理方法能提升性能,但简单的思维链方法在计算开销显著更低的情况下仍具稳健表现。AGORA不仅简化了语言智能体的开发,还通过标准化评估协议为可复现的智能体研究奠定基础。

原文摘要 · Abstract (English)

Language agents powered by large language models (LLMs) have demonstrated remarkable capabilities in understanding, reasoning, and executing complex tasks. However, developing robust agents presents significant challenges: substantial engineering overhead, lack of standardized components, and insufficient evaluation frameworks for fair comparison. We introduce Agent Graph-based Orchestration for Reasoning and Assessment (AGORA), a flexible and extensible framework that addresses these challenges through three key contributions: (1) a modular architecture with a graph-based workflow engine, efficient memory management, and clean component abstraction; (2) a comprehensive suite of reusable agent algorithms implementing state-of-the-art reasoning approaches; and (3) a rigorous evaluation framework enabling systematic comparison across multiple dimensions. Through extensive experiments on mathematical reasoning and multimodal tasks, we evaluate various agent algorithms across different LLMs, revealing important insights about their relative strengths and applicability. Our results demonstrate that while sophisticated reasoning approaches can enhance agent capabilities, simpler methods like Chain-of-Thought often exhibit robust performance with significantly lower computational overhead. AGORA not only simplifies language agent development but also establishes a foundation for reproducible agent research through standardized evaluation protocols.

智能体大模型可复现评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。