arXiv:2605.09005cs.ROcs.AI2026-05被引 2

为视觉语言动作模型设计了隐蔽水印验证系统,防止盗用。

Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models

论文配图:Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 训练时向视觉数据注入秘密信息,嵌入隐蔽水印。
  • 验证时通过触发器与分类器检测水印,准确率超95%。
  • 适用于开源模型防篡改,适合机器人控制领域使用。

视觉语言动作模型(VLAs)通过多模态输入实现端到端决策策略,支持通用机器人控制。随着训练好的VLAs被广泛共享和适配,保护模型所有权对安全部署和负责任的开源使用至关重要。本文提出GuardVLA,首个专为VLAs设计的基于后门的产权验证框架。GuardVLA在训练阶段通过向具身视觉数据中注入秘密消息,嵌入隐蔽且无害的后门水印。发布后,采用交换-检测机制,利用触发器投影器和外部分类器头,根据预测概率激活并检测嵌入的后门。在多个数据集、模型架构及适配设置下的大量实验表明,GuardVLA可实现可靠产权验证,同时保持良性任务性能。进一步结果显示,嵌入的水印在模型发布后适配下仍可检测。

原文摘要 · Abstract (English)

Vision-Language-Action models (VLAs) support generalist robotic control by enabling end-to-end decision policies directly from multi-modal inputs. As trained VLAs are increasingly shared and adapted, protecting model ownership becomes essential for secure deployment and responsible open-source usage. In this paper, we present GuardVLA, the first backdoor-based ownership verification framework specifically designed for VLAs. GuardVLA embeds a stealthy and harmless backdoor watermark into the protected model during training by injecting secret messages into embodied visual data. For post-release verification, we propose a swap-and-detect mechanism, in which the trigger projector and an external classifier head are used to activate and detect the embedded backdoor based on prediction probabilities. Extensive experiments across multiple datasets, model architectures, and adaptation settings demonstrate that GuardVLA enables reliable ownership verification while preserving benign task performance. Further results show that the embedded watermark remains detectable under post-release model adaptation.

模型产权后门水印机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。