提出T-Core框架,从预训练模型和数据中识别可信元素,防御资源受限下的后门攻击。
Secure Transfer Learning: Training Clean Models Against Backdoor in (Both) Pre-trained Encoders and Downstream Datasets
- 通过定位可信数据与神经元,构建主动防御机制。
- 在5种编码器污染、7种数据集污染攻击中均有效,优于14种基线方法。
- 适合需在资源有限下保障模型安全的开发者使用。
基于预训练编码器的迁移学习已成为现代机器学习的核心,能高效适配多样任务。然而,预训练与下游适应的结合扩大了攻击面,使模型在编码器和数据集层面均面临复杂的后门嵌入风险——这一问题在以往研究中常被忽视。此外,用户通常缺乏充足计算资源,导致通用后门防御效果受限,难以与从零训练相媲美。本文研究在资源受限场景下如何缓解潜在后门风险。我们系统分析现有防御策略,发现多数依赖反应式流程,其假设无法扩展至未知威胁、新型攻击或不同训练范式。为此,我们提出主动防御思维,强调识别干净要素,并提出可信核心(T-Core)自举框架,注重定位可信数据与神经元以增强模型安全性。实证评估表明T-Core有效性显著:针对5种编码器污染攻击、7种数据集污染攻击及14种基线防御,在五个基准数据集上覆盖四种场景,应对三类潜在后门威胁。
原文摘要 · Abstract (English)
Transfer learning from pre-trained encoders has become essential in modern machine learning, enabling efficient model adaptation across diverse tasks. However, this combination of pre-training and downstream adaptation creates an expanded attack surface, exposing models to sophisticated backdoor embeddings at both the encoder and dataset levels--an area often overlooked in prior research. Additionally, the limited computational resources typically available to users of pre-trained encoders constrain the effectiveness of generic backdoor defenses compared to end-to-end training from scratch. In this work, we investigate how to mitigate potential backdoor risks in resource-constrained transfer learning scenarios. Specifically, we conduct an exhaustive analysis of existing defense strategies, revealing that many follow a reactive workflow based on assumptions that do not scale to unknown threats, novel attack types, or different training paradigms. In response, we introduce a proactive mindset focused on identifying clean elements and propose the Trusted Core (T-Core) Bootstrapping framework, which emphasizes the importance of pinpointing trustworthy data and neurons to enhance model security. Our empirical evaluations demonstrate the effectiveness and superiority of T-Core, specifically assessing 5 encoder poisoning attacks, 7 dataset poisoning attacks, and 14 baseline defenses across five benchmark datasets, addressing four scenarios of 3 potential backdoor threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。