arXiv:2605.05224cs.LGcs.AI2026-05

提出新方法让隐私数据在预训练模型中无法被学习

Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms

论文配图:Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
图 1 · 摘自论文原文
  • 设计分层欺骗策略,将扰动限制在语义有效空间内
  • 在冻结浅层参数时仍能保持数据不可学习性
  • 适合需要保护隐私的预训练模型应用场景

个人数据在模型训练中的未经授权使用已成为日益严重的隐私威胁。无学习样本(UEs)通过在良性样本中嵌入难以察觉的扰动来阻碍特征学习,从而应对该问题。然而,现有研究主要在从头训练设置下评估UEs,对其在广泛采用的预训练-微调(PF)范式下的行为仍缺乏探索。本文首次系统性地研究了不同训练范式下的无学习样本。分析发现,加载并冻结预训练权重会显著削弱现有UEs方法的效果。我们通过语义过滤机制解释这一现象:尽管UEs倾向于诱导模型过拟合非语义噪声,从而削弱其语义提取能力,但在PF范式下,冻结的浅层保留数据语义,有效过滤掉如无学习噪声等干扰信息。基于此,我们提出层级欺骗策略——浅层语义伪装(SSC),将生成过程限制在语义有效子空间,以绕过预训练权重引入的语义抑制。大量实验表明,该方法在挑战性训练范式(如浅层冻结、语义聚焦预训练)下仍能持续保持数据不可学习性,填补了预训练基础上无学习学习的关键空白。

原文摘要 · Abstract (English)

The unauthorized use of personal data in model training has emerged as a growing privacy threat. Unlearnable examples (UEs) address this issue by embedding imperceptible perturbations into benign examples to obstruct feature learning. However, existing studies mainly evaluate UEs under from-scratch training settings, leaving their behavior under the widely adopted pretraining-finetuning (PF) paradigm largely unexplored. In this work, we provide the first systematic investigation of unlearnable examples across diverse training paradigms. Our analysis reveals that loading and freezing pretrained weights significantly weakens the effectiveness of existing UEs methods. We further explain these findings through semantic filtering: while UEs tend to induce models to overfit non-semantic noise, thereby weakening their semantic extraction capabilities, under the PF paradigm, frozen shallow layers preserve data semantics, effectively filtering out distracting information like unlearnable noise. Guided by these insights, we propose a hierarchical deception strategy, Shallow Semantic Camouflage (SSC), that confines the generation process to a semantically valid subspace, aiming to bypass the semantic suppression introduced by pretrained weights. Extensive experiments demonstrate that our method consistently preserves data unlearnability even under challenging training paradigms, such as shallow-layer freezing and semantic-focused pretraining (SF-Pretrain), bridging the critical gap in pretrain-based unlearnable learning.

隐私保护无学习样本预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。