不修改模型参数,就能让深度模型更抗干扰。
Robustness Reprogramming for Representation Learning
- 提出非线性鲁棒模式匹配新方法,增强特征稳定性。
- 在多种模型上验证,可有效提升对对抗扰动的防御能力。
- 适合需要快速加固现有模型的开发者和研究者。
本文针对表示学习中的一个根本性难题:在不改变已训练好的深度学习模型参数的前提下,能否重新编程以增强其对对抗性或噪声输入扰动的鲁棒性?为此,我们重新审视表示学习中的核心特征变换机制,提出一种新型非线性鲁棒模式匹配技术作为替代方案。此外,设计了三种模型重编程范式,在不同效率需求下灵活控制鲁棒性。在从基础线性模型、MLP到浅层与现代深层卷积网络等多种学习模型上的全面实验和消融研究证明了该方法的有效性。本工作不仅为提升深度学习对抗防御能力开辟了一条新颖且互补的方向,也为设计具有稳健统计特性的更可靠人工智能系统提供了新见解。
原文摘要 · Abstract (English)
This work tackles an intriguing and fundamental open challenge in representation learning: Given a well-trained deep learning model, can it be reprogrammed to enhance its robustness against adversarial or noisy input perturbations without altering its parameters? To explore this, we revisit the core feature transformation mechanism in representation learning and propose a novel non-linear robust pattern matching technique as a robust alternative. Furthermore, we introduce three model reprogramming paradigms to offer flexible control of robustness under different efficiency requirements. Comprehensive experiments and ablation studies across diverse learning models ranging from basic linear model and MLPs to shallow and modern deep ConvNets demonstrate the effectiveness of our approaches. This work not only opens a promising and orthogonal direction for improving adversarial defenses in deep learning beyond existing methods but also provides new insights into designing more resilient AI systems with robust statistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。