arXiv:2501.11067cs.CLcs.AI2025-01被引 26

IntellAgent用多智能体框架自动构建对话AI评估基准,提升测试真实性与诊断精度。

IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems

  • 基于图模型的政策驱动机制,模拟复杂多策略交互场景。
  • 自动生成多样化合成基准,实现细粒度性能诊断与漏洞定位。
  • 开源模块化设计,适合研究者与开发者优化对话系统部署能力。

大型语言模型正演变为具备自主规划与执行能力的任务导向系统,其在对话AI中的应用需应对多轮对话、领域特定API集成及严格策略约束等挑战。然而,传统评估方法难以捕捉真实交互的复杂性与变异性。我们提出IntellAgent,一个可扩展的开源多智能体评估框架,通过结合策略驱动的图建模、真实事件生成与交互式用户-代理模拟,自动化构建多样化的合成基准。该方法克服了静态人工标注基准在粗粒度指标上的局限,提供精细诊断能力。IntellAgent采用基于图的策略模型,刻画策略间关系、概率与复杂性,精准捕捉智能体能力与策略约束间的细微互动。实验表明,该框架能有效识别关键性能短板,为定向优化提供可行建议。其模块化与开源设计支持新领域、策略与API的无缝集成,促进可复现性与社区协作。结果证明,IntellAgent为推动对话AI从研究到部署的跨越提供了有效工具,项目地址:https://github.com/plurai-ai/intellagent。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are transforming artificial intelligence, evolving into task-oriented systems capable of autonomous planning and execution. One of the primary applications of LLMs is conversational AI systems, which must navigate multi-turn dialogues, integrate domain-specific APIs, and adhere to strict policy constraints. However, evaluating these agents remains a significant challenge, as traditional methods fail to capture the complexity and variability of real-world interactions. We introduce IntellAgent, a scalable, open-source multi-agent framework designed to evaluate conversational AI systems comprehensively. IntellAgent automates the creation of diverse, synthetic benchmarks by combining policy-driven graph modeling, realistic event generation, and interactive user-agent simulations. This innovative approach provides fine-grained diagnostics, addressing the limitations of static and manually curated benchmarks with coarse-grained metrics. IntellAgent represents a paradigm shift in evaluating conversational AI. By simulating realistic, multi-policy scenarios across varying levels of complexity, IntellAgent captures the nuanced interplay of agent capabilities and policy constraints. Unlike traditional methods, it employs a graph-based policy model to represent relationships, likelihoods, and complexities of policy interactions, enabling highly detailed diagnostics. IntellAgent also identifies critical performance gaps, offering actionable insights for targeted optimization. Its modular, open-source design supports seamless integration of new domains, policies, and APIs, fostering reproducibility and community collaboration. Our findings demonstrate that IntellAgent serves as an effective framework for advancing conversational AI by addressing challenges in bridging research and deployment. The framework is available at https://github.com/plurai-ai/intellagent

对话AI多智能体评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。