arXiv:2410.10473cs.LGstat.ML2024-10中稿 · NeurIPS被引 2

结构化状态空间模型的隐式偏差可被干净标签数据破坏,导致泛化失败。

The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean Labels

  • 发现干净标签数据会扭曲模型隐式偏差,破坏泛化能力。
  • 实验证明在独立训练和嵌入神经网络中均存在该现象。
  • 对主流的SSM模型提出潜在安全风险,适合关注模型鲁棒性的研究者。

神经网络依赖隐式偏差:梯度下降倾向于以能泛化到未见数据的方式拟合训练数据。近年来,结构化状态空间模型(SSMs)因其高效性成为Transformer的替代方案。已有研究认为,当数据由低维教师模型生成时,SSMs的隐式偏差能实现良好泛化。本文重新审视该设定,首次证明一种此前未被发现的现象:尽管多数训练数据下隐式偏差仍促进泛化,但特定样本的加入会彻底扭曲其偏差,导致泛化完全失效。值得注意的是,这些特殊样本虽由教师模型正确标注(即具有干净标签),仍会造成灾难性影响。我们通过独立训练及作为非线性神经网络组件的实验验证了该现象。在对抗机器学习中,这种以干净标签破坏泛化的攻击称为‘干净标签投毒’。鉴于SSMs广泛应用,明确其对这类攻击的敏感性并发展防御方法,是亟待推进的关键研究方向。

原文摘要 · Abstract (English)

Neural networks are powered by an implicit bias: a tendency of gradient descent to fit training data in a way that generalizes to unseen data. A recent class of neural network models gaining increasing popularity is structured state space models (SSMs), regarded as an efficient alternative to transformers. Prior work argued that the implicit bias of SSMs leads to generalization in a setting where data is generated by a low dimensional teacher. In this paper, we revisit the latter setting, and formally establish a phenomenon entirely undetected by prior work on the implicit bias of SSMs. Namely, we prove that while implicit bias leads to generalization under many choices of training data, there exist special examples whose inclusion in training completely distorts the implicit bias, to a point where generalization fails. This failure occurs despite the special training examples being labeled by the teacher, i.e. having clean labels! We empirically demonstrate the phenomenon, with SSMs trained independently and as part of non-linear neural networks. In the area of adversarial machine learning, disrupting generalization with cleanly labeled training examples is known as clean-label poisoning. Given the proliferation of SSMs, we believe that delineating their susceptibility to clean-label poisoning, and developing methods for overcoming this susceptibility, are critical research directions to pursue.

状态空间模型隐式偏差干净标签攻击泛化失败

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。