arXiv:2605.15520cs.LGcs.AI2026-05

分布式训练中,单个参与者可伪造数据贡献度而不影响模型性能。

On the Fragility of Data Attribution When Learning Is Distributed

  • 通过注入微小合成数据,利用标签非独立同分布特性提升自身贡献度
  • 攻击后模型准确率不变,且能改变其他参与方的贡献排名
  • 揭示了数据归属机制本身存在安全漏洞,适合关注模型治理的研究者

数据归属已成为机器学习流程中定价、审计和治理的关键环节,但现有归属方法隐含假设:归属值能真实反映参与者的贡献。我们证明这一假设可能失效:在标准分布式训练中,单一参与者可通过潜优化注入少量合成数据,在保持全局性能的同时显著夸大自身的归属值。该归属优先攻击利用非独立同分布标签覆盖与评估器敏感性,跨多个数据集、模型及边际效用评估器均成功提升攻击者归属值,并重构良性客户端间的相对归属结构,而不会降低准确率或触发基于几何的防御机制。结果表明,归属本身已成为新的攻击面,亟需构建抗攻击且激励相容的评分机制。

原文摘要 · Abstract (English)

Data attribution has become an important component of pricing, auditing, and governance in machine learning pipelines, yet most attribution methods implicitly assume that attribution values faithfully reflect participants' contributions. We show that this assumption can fail: a single participant in a standard distributed training workflow can substantially inflate its measured attribution value while preserving global utility. Our attribution-first attack uses latent optimization to inject small synthetic batches that preserve utility while exploiting non-IID label coverage and evaluator sensitivities. Across datasets, models, and multiple marginal-utility evaluators, the attack consistently increases the adversary's attribution value and reshapes the relative attribution structure among benign clients without degrading accuracy or triggering geometry-based defenses. These results show that attribution itself forms a new attack surface and motivate the development of attribution-robust and incentive-compatible scoring mechanisms.

数据归属分布式训练安全攻击激励机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。