通过因果审计发现,潜在通信效果常被误判,需控制消息替换来准确评估。
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

- 在接收端边界进行受控消息替换,分离信息来源影响。
- 4B模型中任务内容贡献正向,其他例子消息反而降低性能;8B模型方向反转。
- 建议用此方法评估潜在通信,适合关注多智能体协作可解释性的研究者。
基于大语言模型的多智能体系统中,潜在通信通过连续内部表征传递信息,但高表达能力不等于接收方使用了任务相关信息。仅靠最终任务表现无法判断效果是源于消息存在、特定示例内容,还是另一智能体提供的信息。本文提出一种因果审计方法,在发送方表征进入接收方的边界处实施受控消息替换。四种消息设置支持五项测量:编码的发送方信息、接收方对消息存在与身份的敏感性、示例特定内容的任务价值,以及额外智能体提供的附加值。在GSM8K、ARC-C和MATH-500上对Qwen3-4B与Qwen3-8B的潜在接力任务进行测试。在GSM8K上,4B模型整体性能下降1.00个百分点,分解为其他示例消息保留-6.17点,示例特定内容贡献+5.17点;8B模型中两者方向逆转。在MATH-500上,4B模型提升15.00点,其中8.33点来自其他示例消息,6.67点来自示例特定内容;8B模型则主要依赖前者。自替换对比进一步表明,示例特定内容与跨智能体价值相互独立。结果表明,整体准确率无法揭示潜在消息如何影响接收方,推动受控消息比较作为潜在通信的标准评估方法。
原文摘要 · Abstract (English)
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled message replacements at the boundary where the sender-produced representation enters the receiver. Four message settings support five measurements of encoded sender information, receiver sensitivity to message presence and identity, the task value of example-specific content, and the additional value supplied by a separate agent. We apply the audit to latent relay with Qwen3-4B and Qwen3-8B on GSM8K, ARC-C, and MATH-500. On GSM8K, the Qwen3-4B overall performance effect of -1.00 percentage point decomposes into a -6.17-point effect retained by an other-example message and a +5.17-point effect attributable to example-specific content; both component directions reverse at 8B. On MATH-500, the Qwen3-4B gain of 15.00 points comprises 8.33 points retained by an other-example message and 6.67 points attributable to example-specific content, while the 8B gain is dominated by the former component. Self-substitution comparisons further show that example-specific content and other-agent value are distinct. These results show that aggregate accuracy does not identify how a latent message affects the receiver and motivate controlled message comparisons as a standard evaluation for latent communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。