提出可验证的防御机制,保护视觉语言动作模型免受物理贴片攻击
CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

- 通过校准行为一致动作区域和确定性覆盖掩码实现认证防御
- 在模拟与真实场景中均有效抵御贴片攻击,保证任务成功率
- 适用于对抗自适应攻击,且不依赖贴片内容或生成方式
视觉语言动作(VLA)策略易受局部物理扰动影响,但现有可认证贴片防御仅针对离散标签,无法直接对连续、时序相关动作进行认证。本文提出 CertVLA,一种针对有界贴片与纹理攻击下闭环 VLA 控制的可认证防御方法。CertVLA 构建行为一致的动作区域,并利用确定性覆盖掩码确保至少一个预测无攻击影响。具体而言,通过归一化每对掩码间的动作不一致度,在任一第二掩码下仍保持一致的单掩码锚点才被接受;随后将最大-最小-最大回合得分校准为有限样本下的干净覆盖率。将查询级决策联合扩展至完整闭环推演,实现端到端动作认证。进一步证明,在满足有界支持威胁模型的自适应攻击下,所有由 CertVLA 认证的推演仅执行与擦除攻击后的干净预测一致的动作片段。在双掩码推演正确性的条件下,该一致性证书进一步保证任务成功。证书独立于贴片内容、生成方法及物理变换。仿真与真实世界实验均验证了 CertVLA 在贴片攻击下的实证与可认证有效性,模拟验证亦覆盖纹理攻击。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target discrete labels and cannot directly certify continuous, temporally correlated actions. We introduce CertVLA, a certified defense for closed-loop VLA control under bounded patch and texture attacks. CertVLA proposes a calibrated region of behaviorally consistent actions, while deterministic covering masks ensure that at least one checked prediction is attack-free. Specifically, CertVLA normalizes action disagreement by the benign variation of each mask pair and accepts a single-mask anchor only when it remains consistent under every second mask. It then calibrates the resulting max-min-max episode score to provide finite-sample clean coverage. Conjoining query-level decisions extends the action certificate to the complete closed-loop rollout. Furthermore, we prove that against any adaptive attacker satisfying the bounded-support threat model, every rollout certified by CertVLA executes only action chunks consistent with attack-erased clean predictions. Under dual-mask rollout correctness, this consistency certificate further guarantees task success. The certificate is independent of patch content, generation method, and physical transformation. Experiments in simulation and the real world demonstrate the empirical and certified effectiveness of CertVLA against patch attacks, with additional simulation validation on texture attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。