arXiv:2608.06469cs.CRcs.LG2026-08

提出防御金融联合学习中公平性投毒攻击的新方法,确保恶意客户端权重被严格压制。

Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning

论文配图:Fairis: Fairness-Aware Aggregation with Provable Influence Containment against Fairness Poisoning Attacks in Collaborative Machine Learning
图 1 · 摘自论文原文
  • 通过局部公平差异计算客户端权重,服务器端归一化重加权以抑制恶意行为。
  • 在台湾信用数据集上,恶意客户端权重下降41%至54%,且不牺牲模型精度。
  • 理论证明可抵御合谋攻击,适合高公平性要求的金融场景使用。

金融机构间的联合机器学习需兼顾群体公平性与对恶意操纵的鲁棒性。现有公平性感知聚合方法仍易受公平性投毒攻击:恶意客户端在保持准确率的前提下最大化群体差异,可规避基于准确率的拜占庭防御;在本研究威胁模型下,FairFed的基于差距的加权机制可被观测全局公平分的对手利用。本文提出Fairis,一种服务器端重加权方案,每个客户端更新获得归一化权重 $ω_k = \bar{w}_k / \sum_j \bar{w}_j$,其中未归一化得分 $\bar{w}_k = η- \mathcal{F}_k$,$\mathcal{F}_k \in [0,1]$ 为本地等机会差异,$η> 1$ 为安全参数。我们证明了三个性质:单调权重降低(MWR)、群体参与性与非可操控性,并将MWR扩展至合谋少数群体。结合服务器端范数裁剪,可将攻击者对全局模型的偏移控制在 $ω_0 C$ 内,且随其报告的偏差严格递减。假设诚实评分报告(本文未证明该假设),Fairis是唯一评估中保证所有客户端权重严格为正,且能证明恶意者权重随其偏差单调降低的方法;裁剪后的FairFed虽可能达到更低权重,但无保障,且在台湾信用数据集上会直接置零。对于能规避准确率防御、仅损失0.04准确率的隐蔽攻击,在台湾信用数据集上,Fairis使攻击者权重比无偏控基线降低41%至54%。在常规非独立同分布划分下无明显优势,均匀加权消融实验表明,权重压制效果取决于攻击者得分与诚实均值的偏离程度,若诚实群体已不公平则无法提供保护。

原文摘要 · Abstract (English)

Collaborative machine learning among financial institutions must be both group-fair and robust against deliberate adversarial manipulation. Existing fairness-aware aggregation methods remain formally vulnerable to fairness poisoning: a malicious client maximizing group disparity while preserving accuracy evades accuracy-based Byzantine defenses, and in our threat model FairFed's gap-based weighting can be gamed by an adversary who observes the global fairness score. We present Fairis, a server-side reweighting scheme in which each client's update receives the normalized weight $ω_k = \bar{w}_k / \sum_j \bar{w}_j$ built from the unnormalized score $\bar{w}_k = η- \mathcal{F}_k$, with $\mathcal{F}_k \in [0,1]$ the local Equal Opportunity Difference and $η> 1$ a security parameter. We prove three properties, Monotone Weight Reduction (MWR), Demographic Participation, and Non-Gamesmanship, extend MWR to colluding minority coalitions, and show that combining MWR with server-side norm clipping bounds the adversary's displacement of the global model by $ω_0 C$, strictly decreasing in its own reported disparity. Assuming honest score reporting, an assumption this paper does not discharge, Fairis is the only rule evaluated that guarantees every client strictly positive weight while provably reducing an adversary's weight monotonically in its bias; clipped FairFed can reach a lower weight but guarantees nothing and zeroes a client outright on Taiwan Credit. Against an adversary stealthy enough to evade accuracy-based defenses, within 0.04 accuracy of benign, Fairis cuts its weight by 41 to 54% below a size-blind control on Taiwan. On routine non-IID partitions no rule dominates, and a uniform-weighting ablation shows that containment tracks how far the adversary's score separates from the honest mean, providing none when the honest population is already unfair.

公平性联邦学习安全聚合投毒攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。