arXiv:2601.11924cs.LG2026-01被引 1

研究多人协作中奖问题在对抗干扰下的通信策略与验证机制。

Communication-Corruption Coupling and Verification in Cooperative Multi-Objective Bandits

  • 根据通信方式不同,干扰影响可放大至原预算的1到N倍。
  • 共享原始数据会带来N倍额外误差,仅传推荐则无放大。
  • 验证数据能突破高干扰瓶颈,实现稳定学习效果。

我们研究在对抗性干扰和有限验证条件下,具有向量奖励的协同随机多臂老虎机问题。在每轮T中,N个智能体各自选择一个臂,环境生成干净奖励向量,对手根据全局干扰预算Γ对观测反馈进行扰动。性能以坐标单调递增、L-Lipschitz标量化函数ϕ下的团队遗憾衡量,涵盖线性、切比雪夫及光滑单调效用。主要贡献在于揭示了通信-干扰耦合现象:固定环境侧预算Γ,实际有效干扰水平可在Γ至NΓ之间变化,取决于智能体共享原始样本、充分统计量或仅传递臂推荐。通过协议诱导的多重性函数形式化该关系,并给出依赖于有效干扰的遗憾界。作为推论,共享原始样本会导致N倍附加干扰惩罚,而摘要共享与仅推荐共享均保持未放大的O(Γ)项,并实现集中式速率团队遗憾。进一步建立信息论极限:存在不可避免的Ω(Γ)附加惩罚;当Γ=Θ(NT)的高干扰区间,无清洁信息时子线性遗憾不可能实现。最后,我们刻画了全局验证预算ν如何恢复可学习性:在高干扰区验证是必需的,一旦超过识别阈值即足够,经认证的共享使团队遗憾独立于Γ。

原文摘要 · Abstract (English)

We study cooperative stochastic multi-armed bandits with vector-valued rewards under adversarial corruption and limited verification. In each of $T$ rounds, each of $N$ agents selects an arm, the environment generates a clean reward vector, and an adversary perturbs the observed feedback subject to a global corruption budget $Γ$. Performance is measured by team regret under a coordinate-wise nondecreasing, $L$-Lipschitz scalarization $ϕ$, covering linear, Chebyshev, and smooth monotone utilities. Our main contribution is a communication-corruption coupling: we show that a fixed environment-side budget $Γ$ can translate into an effective corruption level ranging from $Γ$ to $NΓ$, depending on whether agents share raw samples, sufficient statistics, or only arm recommendations. We formalize this via a protocol-induced multiplicity functional and prove regret bounds parameterized by the resulting effective corruption. As corollaries, raw-sample sharing can suffer an $N$-fold larger additive corruption penalty, whereas summary sharing and recommendation-only sharing preserve an unamplified $O(Γ)$ term and achieve centralized-rate team regret. We further establish information-theoretic limits, including an unavoidable additive $Ω(Γ)$ penalty and a high-corruption regime $Γ=Θ(NT)$ where sublinear regret is impossible without clean information. Finally, we characterize how a global budget $ν$ of verified observations restores learnability. That is, verification is necessary in the high-corruption regime, and sufficient once it crosses the identification threshold, with certified sharing enabling the team's regret to become independent of $Γ$.

多智能体强化学习对抗干扰协作优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。