arXiv:2606.28839cs.LG2026-06

提出量化大模型多智能体输出耦合的工具,揭示视觉条件可能夸大性能表现。

The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables

论文配图:The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables
图 1 · 摘自论文原文
  • 构建传播张量框架,通过基准比值度量跨模态、智能体和时间的输出耦合强度。
  • 实验证明视觉条件下的超线性效应(CAF=1.40)在关闭图像扰动后降为亚线性(CAF=0.87)。
  • 适用于评估多智能体系统中真实耦合与设计伪影,适合模型审计与仿真验证。

我们提出传播张量(Contagion Tensor),用于量化大语言模型(LLM)在模态、智能体和时间步之间的输出分布耦合程度。从中推导出耦合放大因子(CAF),一种基于比率的无量纲指标,形式为 CAF = E[T_condition] / E[T_baseline],并提供自助法置信区间。我们实现四种CAF变体,并在完整的2x2x2块正交模拟设计中评估最强版本,进行模态特异性消融。消融显示,看似图像条件下的超线性效应(CAF=1.40)在关闭图像扰动模块后降至亚线性(CAF=0.87),变化-0.53,而文本条件影响不变。补充的实时API实验涵盖两个模型族:DeepSeek-Chat(R=30)和GPT-4o-mini(R=15,真实视觉)。在统一角色下,纯文本通信的CAF≈1.0;多样角色引发收敛(CAF=0.88)。对GPT-4o-mini的模型内比较显示:C3(文本)CAF=1.02 vs. C5(真实视觉,R=30)CAF=1.72 [1.700, 1.733],增量+0.70,验证了模拟中的超线性预测。11个条件中已有5个在真实API上测试,6个尚未验证。贡献在于:(1)一个使输出耦合可量化解构的测量工具;(2)一种可迁移的消融协议,任何模块化多智能体模拟器均可采用以区分真实耦合与设计伪影。

原文摘要 · Abstract (English)

We introduce the Contagion Tensor, a measurement framework for quantifying how large language model (LLM) output distributions couple across modalities, agents, and time steps. From the tensor we derive the Coupling Amplification Factor (CAF), a family of ratio-based metrics sharing the form CAF = E[T_condition] / E[T_baseline], providing unitless, baseline-referenced measurement with bootstrap confidence intervals. We instantiate CAF in four variants and evaluate the strongest in a complete 2x2x2 block-orthogonal simulation design with modality-specific ablation. The ablation reveals that an apparent image-condition super-linear effect (CAF = 1.40) collapses to sub-linear (CAF = 0.87) when the image perturbation module is disabled, a shift of -0.53 with zero effect on text conditions. We supplement with real-API experiments across two model families: DeepSeek-Chat (R=30) and GPT-4o-mini (R=15, real vision). Under uniform personas, text-only communication produces CAF approx 1.0 in both models. Diverse personas drive convergence (CAF = 0.88). A within-model comparison on GPT-4o-mini reveals: C3 (text) CAF = 1.02 vs. C5 (real vision, R=30) CAF = 1.72 [1.700, 1.733], delta = +0.70, validating the simulation's super-linear image-condition prediction. Of 11 conditions, 5 have been tested on real APIs and 6 remain unverified. Our contribution is two-layered: (1) a measurement instrument that makes output-distribution coupling quantitatively falsifiable; and (2) a transferable ablation protocol that any modular multi-agent simulator can adopt to distinguish genuine coupling from design artifacts.

多智能体输出耦合模型审计扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。