arXiv:2605.08612cs.RO2026-05ACL

提出自适应对抗调优框架,有效防御视觉语言动作模型的后门攻击

ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models

论文配图:ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
图 1 · 摘自论文原文
  • 根据攻击者能力动态选择梯度解耦策略,解决训练冲突问题
  • 在5%污染率下实现超80%攻击成功率,隐蔽性极强
  • 首次实现数据投毒中语义级触发的隐式解耦攻击

针对视觉-语言-动作(VLA)模型日益严峻的安全漏洞,本研究聚焦于针对视觉路径的后门攻击。我们发现传统攻击范式失效的核心障碍是‘梯度干扰’——由端到端训练中冲突策略引发的优化失败。为此,提出自适应威胁感知对抗调优(ATAAT)框架,其核心‘威胁-方法自适应映射’机制能根据攻击者能力智能选择最优梯度解耦策略。大量实验表明,ATAAT具备显著优势,在仅5%污染率下实现超过80%的目标攻击成功率(TASR > 80%),同时保持极高隐蔽性。该框架可高效处理复杂的语义级触发器,并首次在数据投毒场景中实现隐式解耦攻击。本工作揭示了VLA模型的关键安全漏洞,为未来防御架构提供了理论与方法支持。

原文摘要 · Abstract (English)

Addressing the escalating security vulnerabilities in Vision-Language-Action (VLA) models, this study investigates backdoor attacks targeting the visual pathway. We identify a core obstacle causing the failure of traditional attack paradigms: "Gradient Interference." This phenomenon represents an optimization failure triggered by conflicting strategies during end-to-end training. To resolve this, we propose an Adaptive Threat-Aware Adversarial Tuning (ATAAT) framework. Through its core "Threat-Method Adaptive Mapping" mechanism, ATAAT intelligently selects the optimal gradient decoupling strategy based on the adversary's capabilities. Extensive experiments demonstrate that ATAAT exhibits significant advantages, achieving a highly robust Targeted Attack Success Rate (TASR > 80%) while maintaining extreme stealthiness with merely a 5% poisoning rate. It efficiently handles complex semantic-level triggers and achieves implicit decoupled attacks in data poisoning scenarios for the first time. This work reveals a critical security vulnerability in VLAs and provides theoretical and methodological support for future defense architectures.

后门攻击视觉语言动作对抗调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。