arXiv:2510.25612cs.AIcs.MA2025-10EMNLP被引 2

首次量化评估智能体工作流中各智能体的影响度。

Counterfactual-based Agent Influence Ranker for Agentic AI Workflows

  • 基于反事实分析,动态评估每个智能体的贡献。
  • 在30个场景230项功能上验证,排名稳定且优于基线。
  • 适用于调试、优化和安全审查,适合开发者与研究者使用。

智能体工作流(AAW)是一种由多个基于大模型的智能体协作完成共同目标的自主系统。其高度自治性、广泛应用及持续增长的兴趣凸显了深入理解其运行机制(包括质量与安全)的必要性。目前尚无方法能评估各智能体对最终输出的影响程度。现有技术多依赖静态结构分析,无法适用于推理阶段。本文提出反事实智能体影响度排序器(CAIR),是首个可评估各智能体对输出影响程度的方法,支持离线与推理时使用。通过反事实分析,实现任务无关的评估。我们在自建的AAW数据集上进行评估,包含30种不同用例和230种功能。结果表明,CAIR能产生一致的排名,显著优于基线方法,并有效提升下游任务的效果与相关性。

原文摘要 · Abstract (English)

An Agentic AI Workflow (AAW), also known as an LLM-based multi-agent system, is an autonomous system that assembles several LLM-based agents to work collaboratively towards a shared goal. The high autonomy, widespread adoption, and growing interest in such AAWs highlight the need for a deeper understanding of their operations, from both quality and security aspects. To this day, there are no existing methods to assess the influence of each agent on the AAW's final output. Adopting techniques from related fields is not feasible since existing methods perform only static structural analysis, which is unsuitable for inference time execution. We present Counterfactual-based Agent Influence Ranker (CAIR) - the first method for assessing the influence level of each agent on the AAW's output and determining which agents are the most influential. By performing counterfactual analysis, CAIR provides a task-agnostic analysis that can be used both offline and at inference time. We evaluate CAIR using an AAWs dataset of our creation, containing 30 different use cases with 230 different functionalities. Our evaluation showed that CAIR produces consistent rankings, outperforms baseline methods, and can easily enhance the effectiveness and relevancy of downstream tasks.

智能体系统影响评估反事实分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。