arXiv:2509.15366cs.AI2025-09被引 2

用动态评估诊断多智能体系统认知缺陷并引导改进

Diagnostics of cognitive failures in multi-agent expert systems using dynamic evaluation protocols and subsequent mutation of the processing context

  • 构建包含专家标注与可控变异数据的诊断框架
  • 发现偏见表述、信息提取漂移等隐性错误
  • 适合需要提升LLM智能体可靠性的研究者使用

神经架构从多层感知机演进到大规模Transformer模型,使语言模型(LLMs)在具备记忆、规划和外部工具使用能力时展现出涌现的代理行为。然而,其固有的随机性和多步决策过程使得传统评估方法难以诊断代理性能。本文提出一种诊断框架,不仅评估专家系统,还能促进专家行为向LLM驱动代理的转移。该框架整合了(1)经筛选的专家标注黄金数据集,(2)通过受控行为变异生成的银数据集,以及(3)基于LLM的代理裁判,可评分并提出针对性改进建议。这些建议被嵌入向量化推荐图中,实现专家干预作为可复用的改进轨迹在多个系统实例间传播。我们在一个多智能体招聘助手系统上验证了该框架,发现其能揭示隐性认知失败——如偏见性表述、信息提取漂移和工具误路由——同时引导代理向专家级推理与风格演进。结果为随机、工具增强型LLM代理中的标准化、可复现专家行为转移奠定了基础,推动评估从静态走向主动系统优化。

原文摘要 · Abstract (English)

The rapid evolution of neural architectures - from multilayer perceptrons to large-scale Transformer-based models - has enabled language models (LLMs) to exhibit emergent agentic behaviours when equipped with memory, planning, and external tool use. However, their inherent stochasticity and multi-step decision processes render classical evaluation methods inadequate for diagnosing agentic performance. This work introduces a diagnostic framework for expert systems that not only evaluates but also facilitates the transfer of expert behaviour into LLM-powered agents. The framework integrates (i) curated golden datasets of expert annotations, (ii) silver datasets generated through controlled behavioural mutation, and (iii) an LLM-based Agent Judge that scores and prescribes targeted improvements. These prescriptions are embedded into a vectorized recommendation map, allowing expert interventions to propagate as reusable improvement trajectories across multiple system instances. We demonstrate the framework on a multi-agent recruiter-assistant system, showing that it uncovers latent cognitive failures - such as biased phrasing, extraction drift, and tool misrouting - while simultaneously steering agents toward expert-level reasoning and style. The results establish a foundation for standardized, reproducible expert behaviour transfer in stochastic, tool-augmented LLM agents, moving beyond static evaluation to active expert system refinement.

多智能体认知诊断行为迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。