arXiv:2605.12280cs.SEcs.AI2026-05被引 1

用多智能体迭代审计发现大型提示系统中的51个缺陷

Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance

论文配图:Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance
图 1 · 摘自论文原文
  • 让多个大模型智能体轮流检查提示规范,逐步发现漏洞
  • 九轮审计共找出51处不一致问题,最后两轮无新问题出现
  • 适合关注AI系统可靠性与提示工程质量管理的研究者

多智能体大语言模型系统的提示规范包含跨文件的数据契约与集成逻辑,但极少接受结构化审查。本研究以生产级七车道系统AEGIS为案例,对其7152行规范进行九轮迭代式智能体审计,共发现51个一致性缺陷(各轮数量分别为15、8、12、2、8、1、4、1、0)。我们提出七类后验分类体系,包含明确编码规则,呈现非单调收敛特征,符合编辑递推与审查范围扩展规律,并采用锁定审计协议。此外,在一个公开的合成小规格样本上进行了两次部分复现:跨四个前沿厂商(OpenAI、Anthropic、Google、xAI)的联合审计(12次追踪)成功检出全部5个预设缺陷;对分层子样本的人工评估显示类别一致性Cohen's κ=0.80,严重性一致性κ=0.46。完整可复现资源随论文提交。

原文摘要 · Abstract (English)

Prompt specifications for multi-agent large language model (LLM) systems carry data contracts and integration logic across interdependent files but are rarely subjected to structured-inspection rigor. We report a single-system case study of iterative, agent-driven auditing applied to AEGIS (Autonomous Engineering Governance and Intelligence System), a seven-lane production pipeline whose 7152-line specification surface was audited across nine rounds, surfacing 51 consistency defects (per-round counts of 15, 8, 12, 2, 8, 1, 4, 1, 0). We present a seven-category post hoc taxonomy with explicit coding rules, non-monotonic convergence consistent with cascading edits and audit-scope expansion, and a locked audit protocol. We further report two partial replications on a public synthetic mini-specification: a cross-LLM panel of four frontier vendors (OpenAI, Anthropic, Google, xAI; 12 traces; multi-vendor union detects all five seeded defects) and an inter-rater reliability check on a stratified subsample (Cohen's $κ$ = 0.80 on category, 0.46 on severity). The full reproducibility bundle accompanies the submission.

提示工程多智能体质量保障LLM审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。