arXiv:2604.00449cs.LGcs.MA2026-04

提出新型抗拜占庭分布式优化方法,提升系统鲁棒性。

Convergence of Byzantine-Resilient Gradient Tracking via Probabilistic Edge Dropout

  • 通过自中心投影与概率边丢弃双重防御机制
  • 在完全隔离下线性收敛,误差由随机梯度方差决定
  • 适合存在恶意节点的去中心化学习场景

我们研究在存在可能发送任意恶意消息的拜占庭代理的网络中进行分布式优化。提出一种名为梯度跟踪概率边丢弃(GT-PD)的随机梯度跟踪方法,可在对抗性通信下保持梯度跟踪的收敛性。该方法结合两种互补的防御层:一种通用的自中心投影,将每个接收到的消息裁剪至接收端半径为τ的球内;另一种完全去中心化的概率丢弃规则,由决策与跟踪通道中的双度量信任评分驱动。此设计在限制恶意扰动的同时保留了双随机混合结构,而该性质常在去中心化鲁棒聚合中丢失。在完全拜占庭隔离(p_b=0)条件下,GT-PD线性收敛至仅由随机梯度方差决定的邻域。对于部分隔离(p_b>0),引入梯度跟踪概率边丢弃与漏桶集成(GT-PD-L),采用漏桶积分器控制持续扰动引起的追踪误差累积,实现线性收敛至由随机方差和裁剪-漏桶比共同决定的有界邻域。进一步证明,在两层丢弃策略且p_h=1时,隔离拜占庭代理不会向诚实共识动态引入额外方差。在MNIST上对符号翻转、ALIE及内积操纵攻击的实验表明,GT-PD-L在隐蔽攻击下性能优于坐标修剪均值最高达4.3个百分点。

原文摘要 · Abstract (English)

We study distributed optimization over networks with Byzantine agents that may send arbitrary adversarial messages. We propose \emph{Gradient Tracking with Probabilistic Edge Dropout} (GT-PD), a stochastic gradient tracking method that preserves the convergence properties of gradient tracking under adversarial communication. GT-PD combines two complementary defense layers: a universal self-centered projection that clips each incoming message to a ball of radius $τ$ around the receiving agent, and a fully decentralized probabilistic dropout rule driven by a dual-metric trust score in the decision and tracking channels. This design bounds adversarial perturbations while preserving the doubly stochastic mixing structure, a property often lost under robust aggregation in decentralized settings. Under complete Byzantine isolation ($p_b=0$), GT-PD converges linearly to a neighborhood determined solely by stochastic gradient variance. For partial isolation ($p_b>0$), we introduce \emph{Gradient Tracking with Probabilistic Edge Dropout and Leaky Integration} (GT-PD-L), which uses a leaky integrator to control the accumulation of tracking errors caused by persistent perturbations and achieves linear convergence to a bounded neighborhood determined by the stochastic variance and the clipping-to-leak ratio. We further show that under two-tier dropout with $p_h=1$, isolating Byzantine agents introduces no additional variance into the honest consensus dynamics. Experiments on MNIST under Sign Flip, ALIE, and Inner Product Manipulation attacks show that GT-PD-L outperforms coordinate-wise trimmed mean by up to 4.3 percentage points under stealth attacks.

分布式优化拜占庭容错梯度跟踪去中心化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。