用张量网络压缩多智能体协作,实现无需重训练的高效分布式控制。
SPIN: Decentralized Swarm Control via Tensorized Policy Coordination
- 将局部群体策略建模为张量链,线性缩放避免组合爆炸。
- 在多目标协同与密集交互场景中性能最优,零样本适配新任务。
- 结合离线映射器与无训练重要性加权,不依赖在线优化。
去中心化多智能体群体协调面临联合动作空间指数级膨胀与高开销训练循环的挑战。本文提出群组策略干扰网络(SPIN)框架,将多智能体通信拓扑建模为压缩张量网络。通过将局部多智能体团簇的联合策略张量分解为开放边界条件矩阵乘积态(MPS)链,SPIN以团簇级收缩替代显式指数级联合动作枚举,在固定局部行为和键维数下实现线性可扩展性。为将原始空间几何与离散代数后端连接,引入解耦框架:基于离线评估的轻量级冻结神经映射器,结合基于Radon-Nikodým导数的确定性零样本重要性重加权滤波器。在单目标追踪、去中心化区域覆盖与结构化多目标协同等不同任务场景中验证原型系统。结果表明,SPIN可在追踪、分散/区域覆盖与结构化多目标协同任务间复用,尤其在多目标协同与密集局部交互场景中优势显著,且无需场景特定的在线优化或再训练。
原文摘要 · Abstract (English)
Decentralized multi-agent swarm coordination remains fundamentally challenged by the combinatorial scaling of joint action spaces and high-overhead training or optimization loops when managing localized group behaviors. To address this problem from a different perspective, this paper introduces the Swarm Policy Interference Network (SPIN) framework, which models multi-agent communication topologies as compressed tensor networks. By factorizing the joint policy tensors of local multi-agent cliques into Open Boundary Condition Matrix Product State (MPS) chains, SPIN replaces explicit exponential joint-action enumeration with clique-level contractions that scale linearly in clique length for fixed local behavior and bond dimensions. To connect raw spatial geometry with this discrete algebraic backend without relying on online training loops, we introduce a decoupled framework combining a lightweight frozen neural mapper evaluated offline with a deterministic zero-shot importance-reweighting filter based on the Radon-Nikodým derivative. We evaluate an executable prototype of this framework within a simulation experiment across distinct task regimes: single-target tracking, decentralized area coverage, and structured multi-goal coordination. The results demonstrate that SPIN functions as a reusable decentralized coordination layer across tracking, dispersion / area coverage, and structured multi-goal coordination, with its strongest gains appearing in multi-goal coordination and dense local interaction regimes, without requiring scenario-specific online optimization or retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。