用双隐状态记忆突破视觉多智能体系统扩展瓶颈
Dual Latent Memory for Visual Multi-agent System
- 采用双隐状态记忆实现智能体间高效协作
- 准确率提升2.7%-5.4%,令牌消耗降低21.3%-44.8%
- 适合追求高扩展性与低通信开销的多智能体系统
视觉多智能体系统(VMAS)虽有望通过智能体间协作提升综合能力,但实证发现存在反直觉的“扩展墙”:增加智能体交互轮次常导致性能下降,且令牌消耗呈指数增长。我们归因于以文本为中心的通信带来的信息瓶颈——将感知与思考轨迹转化为离散自然语言必然造成语义损失。为此,我们提出模型无关的L²-VMAS框架,支持双隐状态记忆下的智能体协作;同时解耦感知与思考过程,动态融合双隐状态记忆;并引入基于熵的主动触发机制,以按需访问记忆取代被动传输。在多种骨干网络、规模和多智能体结构上的实验表明,该方法有效打破“扩展墙”,具备优异可扩展性,平均准确率提升2.7%-5.4%,令牌使用量减少21.3%-44.8%。
原文摘要 · Abstract (English)
While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure to the information bottleneck inherent in text-centric communication, where converting perceptual and thinking trajectories into discrete natural language inevitably induces semantic loss. To this end, we propose \textbf{L}$\mathbf{^{2}}$\textbf{-VMAS}, a novel model-agnostic framework that enables inter-agent collaboration with dual latent memories. Furthermore, we decouple the perception and thinking while dynamically synthesizing dual latent memories. Additionally, we introduce an entropy-driven proactive triggering that replaces passive information transmission with efficient, on-demand memory access. Extensive experiments among backbones, sizes, and multi-agent structures demonstrate that our method effectively breaks the "scaling wall" with superb scalability, improving average accuracy by 2.7-5.4% while reducing token usage by 21.3-44.8%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。