提出自适应对抗调优框架,有效防御视觉语言动作模型的后门攻击
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models

- 根据攻击者能力动态选择梯度解耦策略,解决训练冲突问题
- 在5%污染率下实现超80%攻击成功率,隐蔽性极强
- 首次实现数据投毒中语义级触发的隐式解耦攻击
针对视觉-语言-动作(VLA)模型日益严峻的安全漏洞,本研究聚焦于针对视觉路径的后门攻击。我们发现传统攻击范式失效的核心障碍是‘梯度干扰’——由端到端训练中冲突策略引发的优化失败。为此,提出自适应威胁感知对抗调优(ATAAT)框架,其核心‘威胁-方法自适应映射’机制能根据攻击者能力智能选择最优梯度解耦策略。大量实验表明,ATAAT具备显著优势,在仅5%污染率下实现超过80%的目标攻击成功率(TASR > 80%),同时保持极高隐蔽性。该框架可高效处理复杂的语义级触发器,并首次在数据投毒场景中实现隐式解耦攻击。本工作揭示了VLA模型的关键安全漏洞,为未来防御架构提供了理论与方法支持。
原文摘要 · Abstract (English)
Addressing the escalating security vulnerabilities in Vision-Language-Action (VLA) models, this study investigates backdoor attacks targeting the visual pathway. We identify a core obstacle causing the failure of traditional attack paradigms: "Gradient Interference." This phenomenon represents an optimization failure triggered by conflicting strategies during end-to-end training. To resolve this, we propose an Adaptive Threat-Aware Adversarial Tuning (ATAAT) framework. Through its core "Threat-Method Adaptive Mapping" mechanism, ATAAT intelligently selects the optimal gradient decoupling strategy based on the adversary's capabilities. Extensive experiments demonstrate that ATAAT exhibits significant advantages, achieving a highly robust Targeted Attack Success Rate (TASR > 80%) while maintaining extreme stealthiness with merely a 5% poisoning rate. It efficiently handles complex semantic-level triggers and achieves implicit decoupled attacks in data poisoning scenarios for the first time. This work reveals a critical security vulnerability in VLAs and provides theoretical and methodological support for future defense architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。