发现现有不可学习数据在多任务下仍可被模型利用,提出新指标评估其真实不可学习性。
How Far Are We from True Unlearnability?
- 通过分析损失曲面差异,揭示不可学习数据的优化路径局限性。
- 提出SAL指标量化参数不可学习性,发现仅部分参数路径存在显著差异。
- 设计UD度量方法,验证主流方法在多任务中仍存学习漏洞,适合安全与隐私研究者参考。
高质量数据对大模型至关重要,但未经授权使用数据严重损害数据所有者权益。为此,研究者提出不可学习样本(UEs)以破坏数据训练可用性。理论上,这些数据应跨任务无法提升模型性能。然而,在多任务数据集Taskonomy上,我们发现UEs在语义分割等任务中仍表现良好,未能实现跨任务不可学习性。这引发核心问题:我们距离真正的不可学习还有多远?本文从模型优化角度出发,通过简单架构观察干净与污染模型的收敛差异,发现只有部分关键参数优化路径存在显著区别,表明损失曲面与不可学习性密切相关。基于此,提出尖锐感知可学习性(SAL)来量化参数不可学习性,并构建不可学习距离(UD)衡量清洁与污染模型间SAL分布差异。最终,在主流不可学习方法上进行基准测试,揭示当前方法能力边界,推动社区对不可学习性实际效果的认知。
原文摘要 · Abstract (English)
High-quality data plays an indispensable role in the era of large models, but the use of unauthorized data for model training greatly damages the interests of data owners. To overcome this threat, several unlearnable methods have been proposed, which generate unlearnable examples (UEs) by compromising the training availability of data. Clearly, due to unknown training purposes and the powerful representation learning capabilities of existing models, these data are expected to be unlearnable for models across multiple tasks, i.e., they will not help improve the model's performance. However, unexpectedly, we find that on the multi-task dataset Taskonomy, UEs still perform well in tasks such as semantic segmentation, failing to exhibit cross-task unlearnability. This phenomenon leads us to question: How far are we from attaining truly unlearnable examples? We attempt to answer this question from the perspective of model optimization. To this end, we observe the difference in the convergence process between clean and poisoned models using a simple model architecture. Subsequently, from the loss landscape we find that only a part of the critical parameter optimization paths show significant differences, implying a close relationship between the loss landscape and unlearnability. Consequently, we employ the loss landscape to explain the underlying reasons for UEs and propose Sharpness-Aware Learnability (SAL) to quantify the unlearnability of parameters based on this explanation. Furthermore, we propose an Unlearnable Distance (UD) to measure the unlearnability of data based on the SAL distribution of parameters in clean and poisoned models. Finally, we conduct benchmark tests on mainstream unlearnable methods using the proposed UD, aiming to promote community awareness of the capability boundaries of existing unlearnable methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。