arXiv:2602.07358cs.LG2026-02

让表格数据无法被训练,防止隐私泄露。

UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding

  • 分离高敏感与冗余特征,用隐藏捷径干扰学习
  • 在多个数据集上使非法训练接近随机表现
  • 适合金融医疗等敏感表格数据的隐私保护

不可学习样本(UE)已成为防止未经授权使用私有视觉数据训练模型的有效手段,但将其扩展至表格数据面临挑战。金融与医疗领域的表格数据高度敏感,现有UE方法迁移效果差,因表格特征同时包含数值与类别约束,且显著性稀疏,学习过程受少数维度主导。在谱主导条件下,我们证明当毒物谱强于干净语义谱时,可实现认证不可学习性。基于此,提出无监督表格数据不可学习生成方法UTOPIA,利用特征冗余将优化解耦为两条通道:高显著性特征用于语义混淆,低显著性冗余特征用于嵌入高度相关的隐藏捷径,生成满足约束的主导捷径,同时保持表格数据有效性。在多个表格数据集与模型上的实验表明,UTOPIA使未经授权训练性能趋近随机水平,显著优于强基线,且跨架构迁移良好。

原文摘要 · Abstract (English)

Unlearnable examples (UE) have emerged as a practical mechanism to prevent unauthorized model training on private vision data, while extending this protection to tabular data is nontrivial. Tabular data in finance and healthcare is highly sensitive, yet existing UE methods transfer poorly because tabular features mix numerical and categorical constraints and exhibit saliency sparsity, with learning dominated by a few dimensions. Under a Spectral Dominance condition, we show certified unlearnability is feasible when the poison spectrum overwhelms the clean semantic spectrum. Guided by this, we propose Unlearnable Tabular Data via DecOuPled Shortcut EmbeddIng (UTOPIA), which exploits feature redundancy to decouple optimization into two channels: high saliency features for semantic obfuscation and low saliency redundant features for embedding a hyper correlated shortcut, yielding constraint-aware dominant shortcuts while preserving tabular validity. Extensive experiments across tabular datasets and models show UTOPIA drives unauthorized training toward near random performance, outperforming strong UE baselines and transferring well across architectures.

隐私保护表格数据不可学习对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。