攻击遥感领域迁移学习模型,用神经元操控实现高效越狱
On the Adversarial Vulnerabilities of Transfer Learning in Remote Sensing
- 通过操纵预训练模型的单个或多个脆弱神经元生成可迁移扰动
- 无需领域知识,在多个遥感数据集上均实现高攻击成功率
- 揭示迁移学习安全风险,适合关注模型防御的研究者参考
在遥感领域广泛使用通用计算机视觉任务的预训练模型,能显著降低训练成本并提升性能。然而,这一做法也引入了安全隐患,公开的预训练模型可能被用作代理攻击下游模型。本文提出一种新型对抗性神经元操控方法,通过选择性地操纵预训练模型中的单个或多个神经元生成可迁移扰动。与现有攻击方法不同,该方法无需依赖特定领域信息,具备更强的普适性和效率。通过针对多个脆弱神经元施加扰动,攻击效果更优,暴露出深度学习模型的关键漏洞。在多种模型和遥感数据集上的实验验证了该方法的有效性。这种低访问门槛的对抗性神经元操控技术凸显了迁移学习模型的重大安全风险,强调在面向安全关键型遥感任务时,亟需设计更鲁棒的防御机制。
原文摘要 · Abstract (English)
The use of pretrained models from general computer vision tasks is widespread in remote sensing, significantly reducing training costs and improving performance. However, this practice also introduces vulnerabilities to downstream tasks, where publicly available pretrained models can be used as a proxy to compromise downstream models. This paper presents a novel Adversarial Neuron Manipulation method, which generates transferable perturbations by selectively manipulating single or multiple neurons in pretrained models. Unlike existing attacks, this method eliminates the need for domain-specific information, making it more broadly applicable and efficient. By targeting multiple fragile neurons, the perturbations achieve superior attack performance, revealing critical vulnerabilities in deep learning models. Experiments on diverse models and remote sensing datasets validate the effectiveness of the proposed method. This low-access adversarial neuron manipulation technique highlights a significant security risk in transfer learning models, emphasizing the urgent need for more robust defenses in their design when addressing the safety-critical remote sensing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。