arXiv:2601.19967cs.LGcs.AI2026-01中稿 · ICLR被引 1

用线性模型生成不可学习数据,速度快且效果不输深度模型。

Perturbation-Induced Linearization: Constructing Unlearnable Data with Solely Linear Classifiers

  • 仅用线性模型生成扰动,降低计算开销。
  • 在短时间内实现与深度模型相当的不可学习效果。
  • 揭示了扰动导致模型线性化的机制,解释其有效性。

从网络收集数据训练深度模型已成常态,引发未经授权使用数据的担忧。为缓解此问题,不可学习样本通过引入人眼难以察觉的扰动,使模型无法有效学习。然而,现有方法通常依赖深度神经网络作为代理模型生成扰动,带来显著计算成本。本文提出扰动诱导线性化(PIL),仅使用线性代理模型生成扰动,计算效率大幅提升,同时性能与现有方法相当或更优。我们进一步揭示不可学习样本的关键机制:诱导深度模型线性化,解释了PIL为何能在极短时间内取得良好效果。此外,我们分析了百分比部分扰动下不可学习样本的性质。本工作不仅提供了实用的数据保护方案,还深入揭示了不可学习样本的有效性来源。

原文摘要 · Abstract (English)

Collecting web data to train deep models has become increasingly common, raising concerns about unauthorized data usage. To mitigate this issue, unlearnable examples introduce imperceptible perturbations into data, preventing models from learning effectively. However, existing methods typically rely on deep neural networks as surrogate models for perturbation generation, resulting in significant computational costs. In this work, we propose Perturbation-Induced Linearization (PIL), a computationally efficient yet effective method that generates perturbations using only linear surrogate models. PIL achieves comparable or better performance than existing surrogate-based methods while reducing computational time dramatically. We further reveal a key mechanism underlying unlearnable examples: inducing linearization to deep models, which explains why PIL can achieve competitive results in a very short time. Beyond this, we provide an analysis about the property of unlearnable examples under percentage-based partial perturbation. Our work not only provides a practical approach for data protection but also offers insights into what makes unlearnable examples effective.

数据隐私不可学习数据线性模型扰动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。