arXiv:2602.01834cs.RO2026-02被引 3

通过可解释字典学习,实时识别并抑制视觉语言动作模型中的危险概念。

Concept-Based Dictionary Learning for Inference-Time Safety in Vision Language Action Models

  • 从隐藏激活中学习稀疏可解释的语义字典,定位有害概念方向。
  • 在风险阈值超标时动态削弱危险成分,攻击成功率降低超70%。
  • 无需重训练,适配多种模型,适合部署于真实机器人系统。

视觉语言动作(VLA)模型通过将多模态指令转化为可执行行为,实现感知-行动闭环,但这一能力也放大了安全风险:仅在大语言模型中生成有毒文本的越狱攻击,在具身系统中可能引发危险物理行为。现有防御方法如对齐、过滤或提示加固,干预时机过晚或发生在错误模态,导致融合表示仍可被利用。本文提出一种基于概念的字典学习框架,用于推理时的安全控制。通过从隐藏激活中学习稀疏且可解释的字典,识别有害概念方向,并在风险估计超过阈值时衰减相关成分。在Libero-Harm、BadRobot、RoboPair和IS-Bench上的实验表明,该方法达到当前最优防御性能,攻击成功率下降超过70%,同时保持任务成功率。关键优势在于其即插即用、模型无关特性,无需重新训练,可无缝集成至多种VLA模型。据我们所知,这是首个面向具身系统的推理时概念级安全方法,推动了VLA模型的可解释性与安全部署。

原文摘要 · Abstract (English)

Vision Language Action (VLA) models close the perception action loop by translating multimodal instructions into executable behaviors, but this very capability magnifies safety risks: jailbreaks that merely yield toxic text in LLMs can trigger unsafe physical actions in embodied systems. Existing defenses alignment, filtering, or prompt hardening intervene too late or at the wrong modality, leaving fused representations exploitable. We introduce a concept based dictionary learning framework for inference time safety control. By learning sparse, interpretable dictionaries from hidden activations, our method identifies harmful concept directions and attenuates risky components when the estimated risk exceeds a threshold. Experiments on Libero-Harm, BadRobot, RoboPair, and IS-Bench show that our approach achieves state-of-the-art defense performance, cutting attack success rates by over 70\% while maintaining task success. Crucially, the framework is plug-in and model-agnostic, requiring no retraining and integrating seamlessly with diverse VLAs. To our knowledge, this is the first inference time concept based safety method for embodied systems, advancing both interpretability and safe deployment of VLA models.

安全控制可解释性具身智能字典学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。