arXiv:2605.06696cs.AIcs.LG2026-05

通过分析神经网络内部表示,用谱聚类识别多智能体系统中的隐性联盟。

Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

论文配图:Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
图 1 · 摘自论文原文
  • 基于智能体隐藏状态构建互信息图,再用谱聚类划分联盟边界。
  • 在强化学习和大语言模型中均成功识别出动态联盟结构,准确排除虚假协同信号。
  • 适合关注AI安全、对齐与分布式系统涌现行为的研究者使用。

多个交互式AI智能体可能形成联盟,产生影响AI安全与对齐的群体级组织。然而,仅观察行为难以区分真实信息耦合与虚假相似性,因为关键联盟可能在行为变化前就已存在于内部表示层面。本文提出一种从多智能体系统内部神经表示中检测联盟结构的实用方法:构建智能体间隐藏状态的成对互信息图,并应用谱聚类以识别最显著的联盟边界。我们在两个领域验证该方法:首先在多智能体强化学习环境中,成功恢复预设的层次化与动态联盟结构,并正确排除无信息耦合的行为协调假阳性;其次在大型语言模型中,识别出描述性提示所暗示的联盟结构,追踪动态团队重组,并揭示显式标签主导于冲突交互模式的表征层级。两种场景下,所得划分均展现出标量跨智能体互信息无法分辨的子群组织。结果表明,通过谱聚类分析隐藏状态互信息可提供一种可扩展的诊断工具,用于监测分布式AI系统中的涌现结构。

原文摘要 · Abstract (English)

Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. However, observing agent behavior alone is often insufficient to distinguish genuine informational coupling from spurious similarity, as consequential coalitions may form at the level of internal representations before any overt behavioral change is apparent. Here, we introduce a practical method for detecting coalition structure from the internal neural representations of multi-agent systems. The approach constructs a pairwise mutual-information graph from the hidden states of agents and applies spectral partitioning to identify the most salient coalition boundary. We validate this method in two domains. First, in multi-agent reinforcement learning environments, the method successfully recovers programmed hierarchical and dynamic coalition structures and correctly rejects false positives arising from behavioral coordination without informational coupling. Second, using a large language model, the method identifies coalition structures implied by descriptive prompts, tracks dynamic team reassignments, and reveals a representational hierarchy where explicit labels dominate over conflicting interaction patterns. Across both settings, the recovered partition reveals subgroup organization that a scalar cross-agent mutual-information measure cannot distinguish. The results demonstrate that analyzing hidden-state mutual information through spectral partitioning provides a scalable diagnostic for identifying representational coalitions, offering a valuable tool for monitoring emergent structure in distributed AI systems.

多智能体联盟检测表征分析谱聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。