arXiv:2510.04303cs.MAcs.AI2025-10被引 7

提出可验证的多智能体共谋检测方法,解决信任漏洞与审计难问题。

Audit the Whisper: Detecting Steganographic Collusion in Multi-Agent LLMs

  • 基于信道容量理论量化干预措施对共谋能力的抑制效果
  • 构建覆盖多种场景的共谋基准测试集,支持可复现审计
  • 融合多指标检测机制,误报率低于千分之一,适合高可靠性场景

大型语言模型的多智能体部署正广泛应用于市场、分配与治理流程,但智能体间的隐秘协作可能悄然破坏信任与社会福利。现有审计方法依赖启发式规则,缺乏理论保证,难以跨任务迁移,且缺乏独立复现所需基础设施。本文提出 Audit the Whisper,一个具备理论、基准设计、检测与可复现性的会议级研究成果:(i) 通过信道容量分析证明改写、限速、角色轮换等干预手段会带来可量化的容量损耗,利用配对运行的KL散度诊断实现有限样本下的互信息阈值收紧,并提供完整证明;(ii) 构建 ColludeBench-v0 基准,涵盖定价、第一价格拍卖、同行评审及 Gemini/Groq API 接口,支持可配置的隐蔽协作方案、确定性清单与奖励追踪;(iii) 设计校准审计流水线,融合跨运行互信息、排列不变性、水印方差与公平性感知接受偏差,每项均控制在 10⁻³ 的假阳性预算内,并通过 1 万次诚实运行和 e-value 鞣马检验验证。在 ColludeBench 及 Secret Collusion、CASE、Perfect Collusion Benchmark、SentinelAgent 等外部套件上的联合元测试中,该方法在固定假阳性率下达到当前最优检测性能;消融实验揭示了审计成本权衡,也发现了仅靠互信息无法察觉的公平性驱动共谋者。我们开源再生脚本、匿名清单与文档,使外部审计员可复现全部图表,满足双盲要求,并以最小代价扩展框架。

原文摘要 · Abstract (English)

Multi-agent deployments of large language models (LLMs) are increasingly embedded in market, allocation, and governance workflows, yet covert coordination among agents can silently erode trust and social welfare. Existing audits are dominated by heuristics that lack theoretical guarantees, struggle to transfer across tasks, and seldom ship with the infrastructure needed for independent replication. We introduce Audit the Whisper, a conference-grade research artifact that spans theory, benchmark design, detection, and reproducibility. Our contributions are: (i) a channel-capacity analysis showing how interventions such as paraphrase, rate limiting, and role permutation impose quantifiable capacity penalties-operationalised via paired-run Kullback--Leibler diagnostics-that tighten mutual-information thresholds with finite-sample guarantees and full proofs; (ii) ColludeBench-v0, covering pricing, first-price auctions, peer review, and hosted Gemini/Groq APIs with configurable covert schemes, deterministic manifests, and reward instrumentation; and (iii) a calibrated auditing pipeline that fuses cross-run mutual information, permutation invariance, watermark variance, and fairness-aware acceptance bias, each tuned to a $10^{-3}$ false-positive budget and validated by 10k honest runs plus an e-value martingale. Across ColludeBench and external suites including Secret Collusion, CASE, Perfect Collusion Benchmark, and SentinelAgent, the union meta-test attains state-of-the-art power at fixed FPR while ablations surface price-of-auditing trade-offs and fairness-driven colluders invisible to MI alone. We release regeneration scripts, anonymized manifests, and documentation so that external auditors can reproduce every figure, satisfy double-blind requirements, and extend the framework with minimal effort.

多智能体共谋检测可复现审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。