模型越复杂,越容易被针对特定群体的隐蔽攻击破坏。
Fragile Giants: Understanding the Susceptibility of Models to Subpopulation Attacks
- 提出理论框架,解释大模型为何会记忆并误判目标子群体。
- 实验显示参数越多的模型,越易受子群体攻击影响。
- 小而可解释的子群体攻击常被模型忽略,威胁隐蔽性强。
随着机器学习模型日益复杂,其鲁棒性与可信度问题愈发突出。数据投毒攻击是关键隐患,其中子群体投毒攻击尤为隐蔽——仅针对数据集中特定子群体进行篡改,整体性能却基本不变。这类攻击在真实场景中可能专门损害少数或代表性不足群体。本文研究模型复杂度对子群体投毒攻击的敏感性,提出理论框架解释:过参数化模型因容量过大,会无意中记忆并错误分类目标子群体。我们在大规模图像与文本数据集上使用主流模型架构进行了广泛实验,结果表明参数越多的模型,越容易遭受子群体投毒攻击。此外,针对更小、人类可解释的子群体的攻击往往难以被检测。这些发现强调了需发展专门防御机制来应对子群体脆弱性。
原文摘要 · Abstract (English)
As machine learning models become increasingly complex, concerns about their robustness and trustworthiness have become more pressing. A critical vulnerability of these models is data poisoning attacks, where adversaries deliberately alter training data to degrade model performance. One particularly stealthy form of these attacks is subpopulation poisoning, which targets distinct subgroups within a dataset while leaving overall performance largely intact. The ability of these attacks to generalize within subpopulations poses a significant risk in real-world settings, as they can be exploited to harm marginalized or underrepresented groups within the dataset. In this work, we investigate how model complexity influences susceptibility to subpopulation poisoning attacks. We introduce a theoretical framework that explains how overparameterized models, due to their large capacity, can inadvertently memorize and misclassify targeted subpopulations. To validate our theory, we conduct extensive experiments on large-scale image and text datasets using popular model architectures. Our results show a clear trend: models with more parameters are significantly more vulnerable to subpopulation poisoning. Moreover, we find that attacks on smaller, human-interpretable subgroups often go undetected by these models. These results highlight the need to develop defenses that specifically address subpopulation vulnerabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。