通过重建视觉标记消除机器人策略中的后门攻击
When Attention Betrays: Erasing Backdoor Attacks in Robotic Policies by Reconstructing Visual Tokens
- 发现后门攻击会劫持深层注意力,形成靠近正常数据的紧凑聚类
- 测试时检测异常注意力区域,用深层线索遮蔽可疑区域并重建无触发图像
- 无需重训练模型,适用于多种机器人平台和任务
视觉-语言-动作(VLA)模型的下游微调提升了机器人性能,但也带来了后门攻击风险。攻击者可在污染数据上预训练VLA,植入隐蔽后门,在推理时触发有害行为。现有防御要么缺乏对多模态后门的机理理解,要么需全模型重训练带来高昂计算成本。本文揭示深层注意力劫持机制:后门会引导晚期注意力,形成靠近干净数据流形的紧凑嵌入聚类。基于此,提出Bera框架——在测试时通过潜在空间定位检测异常注意力,利用深层线索遮蔽可疑区域,并重建无触发图像,打破触发-危险动作映射,同时恢复正确行为。Bera无需重训练或修改训练流程。跨多个具身平台和任务的实验表明,Bera有效维持原始性能,显著降低攻击成功率,并一致恢复良性输出,为机器人系统提供鲁棒且实用的防御方案。
原文摘要 · Abstract (English)
Downstream fine-tuning of vision-language-action (VLA) models enhances robotics, yet exposes the pipeline to backdoor risks. Attackers can pretrain VLAs on poisoned data to implant backdoors that remain stealthy but can trigger harmful behavior during inference. However, existing defenses either lack mechanistic insight into multimodal backdoors or impose prohibitive computational costs via full-model retraining. To this end, we uncover a deep-layer attention grabbing mechanism: backdoors redirect late-stage attention and form compact embedding clusters near the clean manifold. Leveraging this insight, we introduce Bera, a test-time backdoor erasure framework that detects tokens with anomalous attention via latent-space localization, masks suspicious regions using deep-layer cues, and reconstructs a trigger-free image to break the trigger-unsafe-action mapping while restoring correct behavior. Unlike prior defenses, Bera requires neither retraining of VLAs nor any changes to the training pipeline. Extensive experiments across multiple embodied platforms and tasks show that Bera effectively maintains nominal performance, significantly reduces attack success rates, and consistently restores benign behavior from backdoored outputs, thereby offering a robust and practical defense mechanism for securing robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。