群组鲁棒性方法可能放大数据投毒,导致安全与公平的冲突。
Like Oil and Water: Group Robustness Methods and Poisoning Defenses May Be at Odds
- 用启发式方法识别少数群体样本时,会误把投毒样本当正常样本放大。
- 放大投毒样本使攻击成功率从0%升至97%以上,损害模型安全。
- 该问题适用于多数鲁棒性与防御机制,适合关注模型公平与安全的研究者。
群组鲁棒性已成为机器学习中的关键问题,因传统训练范式在少数群体上表现差。缺乏显式群组标注时,现有方法依赖启发式策略识别并放大少数样本。本文首次揭示其核心缺陷:无法区分真实少数样本与投毒样本。放大过程同时提升了投毒样本的影响,使攻击成功率从无放大时的0%增至97%以上。我们还通过不可能性结果证明,在特定假设下标准启发式方法无法实现区分。进一步分析集中式与联邦学习中的最新投毒防御,发现它们同样依赖类似启发式剔除可疑样本,但将少数样本一同清除,导致群组鲁棒性下降——从55%降至41%。由于两者目标相反却使用相同策略,组合应用无法缓解矛盾。本研究揭示了基准驱动的机器学习研究如何掩盖不同指标间的潜在权衡,可能带来严重后果。
原文摘要 · Abstract (English)
Group robustness has become a major concern in machine learning (ML) as conventional training paradigms were found to produce high error on minority groups. Without explicit group annotations, proposed solutions rely on heuristics that aim to identify and then amplify the minority samples during training. In our work, we first uncover a critical shortcoming of these methods: an inability to distinguish legitimate minority samples from poison samples in the training set. By amplifying poison samples as well, group robustness methods inadvertently boost the success rate of an adversary -- e.g., from $0\%$ without amplification to over $97\%$ with it. Notably, we supplement our empirical evidence with an impossibility result proving this inability of a standard heuristic under some assumptions. Moreover, scrutinizing recent poisoning defenses both in centralized and federated learning, we observe that they rely on similar heuristics to identify which samples should be eliminated as poisons. In consequence, minority samples are eliminated along with poisons, which damages group robustness -- e.g., from $55\%$ without the removal of the minority samples to $41\%$ with it. Finally, as they pursue opposing goals using similar heuristics, our attempt to alleviate the trade-off by combining group robustness methods and poisoning defenses falls short. By exposing this tension, we also hope to highlight how benchmark-driven ML scholarship can obscure the trade-offs among different metrics with potentially detrimental consequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。