通过移除法量化多智能体贡献,低成本优化系统性能。
Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution

- 将智能体归因建模为合作博弈,用移除策略评估个体作用。
- 移除低贡献智能体可提效17%、降本35%,且保持诊断准确率。
- 适用于医疗等高风险场景,能解耦诊断与伦理表现并针对性优化。
随着多智能体系统(MAS)日益复杂,识别单个智能体的贡献对系统优化至关重要。现有方法缺乏严谨统一的信用分配框架。本文将智能体归因形式化为合作博弈,参数包括联盟分布、移除协议和目标指标。基于此框架,我们发现留一法(LOO)在计算成本极低的情况下,与组合方法同样有效识别瓶颈智能体。移除协议引发不同博弈:智能体消融可定位结构瓶颈,而内部大模型判断无法忠实反映该行为。此外,为评估特定智能体架构价值,我们提出模型替换归因法。通过替换低贡献智能体的底层模型,在三个基准上实现最高17%的性能提升和最高35%的成本降低。最后,我们将框架应用于医疗多智能体系统,发现诊断准确率与伦理行为贡献常解耦;干预无效角色后,伦理对齐度提升,诊断准确率不变。本工作提供了一种成本可控、原理清晰的多智能体归因与干预方法。
原文摘要 · Abstract (English)
As multi-agent systems (MAS) become increasingly complex, identifying the contributions of individual agents is critical for system optimization. However, existing approaches lack a rigorous, unified framework for credit assignment. In this work, we formalize agent attribution as a cooperative game, parameterized by the coalition distribution, removal protocol, and target metric. Using this framework, we show that Leave-One-Out (LOO) identifies bottleneck agents as effectively as combinatorial methods, but at a fraction of the computational cost. We also demonstrate that removal protocols induce distinct games: Agent ablation isolates structural bottlenecks, whereas introspective LLM judges fail to faithfully approximate this behavior. Furthermore, to evaluate the utility of specific agent backbones, we introduce attribution via model replacement. By substituting underlying models of low-contribution agents, we improve task performance by up to 17% while reducing cost by up to 35% across three benchmarks. Finally, we apply our framework to audit a medical MAS, revealing that agent contributions to diagnostic accuracy and ethical behavior are often decoupled. By intervening on counterproductive roles, we observe an increase in ethics alignment while maintaining diagnostic accuracy. Overall, this work provides a principled approach for cost-effective MAS attribution and intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。