arXiv:2603.04731cs.LG2026-03被引 4

预训练模型会暴露隐蔽扰动数据的虚假线索,新方法通过绑定错误标签来防御。

When Priors Backfire: On the Vulnerability of Unlearnable Examples to Pretraining

  • 设计双层优化机制,强制模型依赖扰动而非语义特征
  • 在多个基准上验证可维持数据不可学习性,对抗预训练先验
  • 适合关注隐私保护与对抗样本防御的研究者

未学习样本(UEs)是一种数据保护策略,通过生成难以察觉的扰动,诱导模型学习虚假关联而非真实语义。本文揭示了一个根本性漏洞:当模型从预训练起点开始学习时,即使数据经过精心构造的扰动保护,预训练先验仍能提供丰富的语义表征,使模型绕过UE引入的捷径,捕获真实特征,从而破坏不可学习性。为此,我们提出BAIT(绑定人工扰动至错误目标),一种新型双层优化框架。内层旨在将扰动样本与真实标签对齐,模拟标准数据-标签关系;外层则主动破坏此对齐,通过强制扰动-错误标签绑定,将样本映射到指定错误目标。该机制有效覆盖了先验的语义引导,迫使模型依赖注入的扰动,进而阻止真实语义的学习。在多个标准基准和预训练主干网络上的广泛实验表明,BAIT能有效缓解预训练先验的影响,维持数据的不可学习性。

原文摘要 · Abstract (English)

Unlearnable Examples (UEs) serve as a data protection strategy that generates imperceptible perturbations to mislead models into learning spurious correlations instead of underlying semantics. In this paper, we uncover a fundamental vulnerability of UEs that emerges when learning starts from a pretrained model. Crucially, our empirical analysis shows that even when data are protected by carefully crafted perturbations, pretraining priors still furnish rich semantic representations that allow the model to circumvent the shortcuts introduced by UEs and capture genuine features, thereby nullifying unlearnability. To address this, we propose BAIT (Binding Artificial perturbations to Incorrect Targets), a novel bi-level optimization formulation. Specifically, the inner level aims at associating the perturbed samples with real labels to simulate standard data-label alignment, while the outer level actively disrupts this alignment by enforcing a mislabel-perturbation binding that maps samples to designated incorrect targets. This mechanism effectively overrides the semantic guidance of priors, forcing the model to rely on the injected perturbations and consequently preventing the acquisition of true semantics. Extensive experiments on standard benchmarks and multiple pretrained backbones demonstrate that BAIT effectively mitigates the influence of pretraining priors and maintains data unlearnability.

数据隐私对抗防御预训练模型不可学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。