提出数据高估攻击与诚实估值方法,解决联邦学习中数据贡献虚报问题。
Data Overvaluation Attack and Truthful Data Valuation in Federated Learning
- 设计数据高估攻击,让恶意客户端夸大自身数据价值。
- 实验表明现有方法易被攻破,新方法在多种场景下保持稳健。
- 适合关注联邦学习激励机制与数据安全的研究者。
在协同机器学习中,数据估值——即评估每个客户端数据对模型的贡献——已成为激励和筛选优质数据的关键任务。然而,现有研究常假设客户端会如实参与估值,忽略了其夸大贡献的实际动机。为揭示这一威胁,本文提出了数据高估攻击,使策略性客户端在联邦学习(一种广泛采用的去中心化协同学习范式)中显著高估自身数据价值。此外,我们提出一种基于贝叶斯的诚实数据估值度量方法,名为 Truth-Shapley。Truth-Shapley 是唯一满足特定理想公理的数据估值方法,在特定条件下确保客户端最优策略是真实估值。实验表明,现有估值方法易受该攻击影响,而 Truth-Shapley 具有鲁棒性和有效性。
原文摘要 · Abstract (English)
In collaborative machine learning (CML), data valuation, i.e., evaluating the contribution of each client's data to the machine learning model, has become a critical task for incentivizing and selecting positive data contributions. However, existing studies often assume that clients engage in data valuation truthfully, overlooking the practical motivation for clients to exaggerate their contributions. To unlock this threat, this paper introduces the data overvaluation attack, enabling strategic clients to have their data significantly overvalued in federated learning, a widely adopted paradigm for decentralized CML. Furthermore, we propose a Bayesian truthful data valuation metric, named Truth-Shapley. Truth-Shapley is the unique metric that guarantees some promising axioms for data valuation while ensuring that clients' optimal strategy is to perform truthful data valuation under certain conditions. Our experiments demonstrate the vulnerability of existing data valuation metrics to the proposed attack and validate the robustness and effectiveness of Truth-Shapley.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。