提出动态任务分解与工具整合框架,提升智能体系统响应能力
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
- 构建可自适应任务图的智能体框架,支持多跳查询与实时调整
- 引入节点F1、结构相似度等新指标,精准评估系统表现
- 基于AsyncHow数据集验证,适用于复杂任务场景的系统优化
大型语言模型的进步正推动自主智能体系统的演进,实现上下文感知的任务动态分解与自动工具选择。本文提出三项核心贡献:一是构建可处理多跳查询、生成并执行任务图、选择合适工具且能应对实时变化的智能体框架;二是引入节点F1分数、结构相似性指数(SSI)和工具F1分数,全面评估智能体系统性能;三是开发基于AsyncHow的数据集,用于分析不同任务复杂度下的智能体行为。研究发现,异步与动态任务图分解显著提升系统响应速度与可扩展性,尤其在复杂多步任务中表现突出。分析表明,结构与节点级指标对顺序任务至关重要,而工具相关指标则主导并行任务的表现。具体而言,结构相似性指数(SSI)是顺序任务性能的最佳预测因子,工具F1分数对并行任务不可或缺。这些发现强调需建立兼顾结构与操作维度的均衡评估方法。所提出的评估框架经实证分析与统计检验验证,为提升智能体在动态环境中的适应性与可靠性提供关键洞见。
原文摘要 · Abstract (English)
Advancements in Large Language Models (LLMs) are revolutionizing the development of autonomous agentic systems by enabling dynamic, context-aware task decomposition and automated tool selection. These sophisticated systems possess significant automation potential across various industries, managing complex tasks, interacting with external systems to enhance knowledge, and executing actions independently. This paper presents three primary contributions to advance this field: - Advanced Agentic Framework: A system that handles multi-hop queries, generates and executes task graphs, selects appropriate tools, and adapts to real-time changes. - Novel Evaluation Metrics: Introduction of Node F1 Score, Structural Similarity Index (SSI), and Tool F1 Score to comprehensively assess agentic systems. - Specialized Dataset: Development of an AsyncHow-based dataset for analyzing agent behavior across different task complexities. Our findings reveal that asynchronous and dynamic task graph decomposition significantly enhances system responsiveness and scalability, particularly for complex, multi-step tasks. Detailed analysis shows that structural and node-level metrics are crucial for sequential tasks, while tool-related metrics are more important for parallel tasks. Specifically, the Structural Similarity Index (SSI) is the most significant predictor of performance in sequential tasks, and the Tool F1 Score is essential for parallel tasks. These insights highlight the need for balanced evaluation methods that capture both structural and operational dimensions of agentic systems. Additionally, our evaluation framework, validated through empirical analysis and statistical testing, provides valuable insights for improving the adaptability and reliability of agentic systems in dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。