arXiv:2608.06520cs.LG2026-08

研究多智能体系统在隐蔽拜占庭攻击下的在线安全学习,提出可保证性能的鲁棒算法。

Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks

  • 基于攻击者是否观测计划动作,构建两种鲁棒MDP模型
  • 证明安全损失由数据生成响应与累计响应差距决定,不可避免
  • 设计分阶段鲁棒学习算法,实现近最优后悔界,适合高安全需求场景

我们研究在拜占庭攻击下的多智能体系统在线协同控制问题。未知且固定的智能体子集被攻陷,可在观察到团队计划联合动作后隐蔽篡改自身坐标。学习者仅能观测计划动作、公开奖励和公开状态,无法获知篡改行为或实际执行的动作。目标是保障安全:优化团队在最坏篡改情况下的性能,并达到最优安全值。我们首先发现攻击者信息决定了几何结构:观测计划动作的攻击者诱导出精确的 $(s,a)$-矩形鲁棒马尔可夫决策过程(MDP),其行是篡改引发的公共结果分布的凸包;盲攻击者则诱导出 $s$-矩形模型。随后,我们识别出安全学习的信息论极限,证明安全后悔分解为返回后悔与累计响应差距 $D_K$ 之和。两个不可区分的一步实例强制期望安全后悔为 $Ω(K)$,而返回后悔为零,表明对 $D_K$ 的依赖不可避免。最后,我们提出一种阶段绑定的鲁棒估计-决策学习器,并证明后悔界为 $ ilde{oldsymbol{O}}ig(H^2S oot{AK}ig) + oldsymbol{E}[D_K]$。本研究为拜占庭攻击下可靠多智能体系统提供了全面的理论与算法基础。

原文摘要 · Abstract (English)

We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint action after observing that plan. The learner observes planned actions, public rewards, and public states, but neither the overwrite nor the executed joint action. Our objective is security: to optimize the team performance against the worst overwrites and achieve the optimal security value. We first show that the attacker's information determines the geometry. An attacker that observes the planned action induces an exact $(s,a)$-rectangular robust Markov decision process (MDP) whose rows are convex hulls of overwrite-induced public-outcome laws, whereas a blind attacker induces an $s$-rectangular model. We then identify the information-theoretic limit of security learning, showing that the security regret decomposes exactly into return regret against the response generating the data and a cumulative response gap $D_K$. Two indistinguishable horizon-one instances force $Ω(K)$ expected security regret while return regret is zero, showing that dependence on $D_K$ is unavoidable. Finally, we develop a stage-tied robust estimation-to-decisions learner and prove a regret bound of $\widetilde{\mathcal O}\!\left(H^2S\sqrt{AK}\right)+\mathbb E[D_K]$. Our studies thus provide comprehensive theoretical and algorithmic foundations of reliable multi-agent systems under Byzantine attacks.

多智能体拜占庭攻击安全学习鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。