在协作任务中,如何用最少通信量学会团队规模阈值?
Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach

- 设计了基于事件触发的去中心化算法,仅在信念变化时通信
- 通信量降低23倍,同时保持与集中式算法相当的性能
- 适用于资源受限的多智能体系统,如分布式机器人协作
在许多多智能体应用中,只有当协作团队规模达到未知阈值时才会获得奖励,否则反馈完全被掩盖。这种掩盖导致可识别性问题:智能体无法区分随机失败与协调不足。我们将其形式化为阈值激活的协作多臂老虎机(TAC-MAB),并分析了集中式与去中心式协调下的表现。集中式算法(C-TAC)的累积遗憾为O(log T),分解为结构搜索项(处理遮蔽反馈下的可行性)和统计监控项(价值估计)。随后提出去中心化的事件触发协议D-TAC,智能体仅在结构信念变化时同步。实验表明,D-TAC相比集中式基线通信量减少23倍,且在保守信念融合下仍保持可行性对齐。结果揭示了在遮蔽反馈下学习的协调成本,并证明无需持续同步即可实现接近集中式的通信效率。
原文摘要 · Abstract (English)
In many multi-agent applications, tasks yield rewards only when executed by a coalition meeting an unknown size threshold; otherwise, feedback is fully censored. This censorship creates an identifiability problem: agents cannot distinguish stochastic failure from insufficient coordination. We formalize this setting as the Threshold-Activated Cooperative Multi-Armed Bandit (TAC-MAB) and analyze it under both centralized and decentralized coordination. We show that a centralized algorithm (C-TAC) achieves cumulative regret O(log T), decomposed into a structural-search term that captures the cost of resolving feasibility under censored feedback and a statistical-monitoring term for value estimation. We then introduce D-TAC, a decentralized event-triggered protocol in which agents synchronize only when their structural beliefs change. Empirically, D-TAC achieves a 23x reduction in communication relative to the centralized baseline while preserving feasibility alignment under conservative belief fusion. These results characterize the coordination cost of learning under censored feedback and show that near-centralized communication efficiency is achievable without continuous synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。