arXiv:2509.23281cs.RO2025-09被引 1

用多模态适配提升机器人防越狱能力,准确率接近100%。

Preventing Robotic Jailbreaking via Multimodal Domain Adaptation

  • 融合文本与视觉特征,通过注意力机制捕捉意图和环境信息
  • 在自动驾驶、海上机器人等场景中检测准确率达近100%
  • 轻量设计适合安全关键的机器人应用,部署开销极小

大型语言模型(LLMs)和视觉-语言模型(VLMs)在机器人环境中日益普及,但易受越狱攻击,可能绕过安全机制并引发真实世界中的不安全或物理伤害行为。基于数据的防御方法如越狱分类器虽有潜力,但在缺乏专用数据集的领域中泛化能力差,限制了其在机器人等安全关键场景的应用。为此,我们提出 J-DAPT,一种基于注意力融合与域适应的轻量级多模态越狱检测框架。该框架整合文本与视觉嵌入,同时将通用越狱数据集与特定领域的参考数据对齐。在自动驾驶、海洋机器人和四足导航任务中的评估显示,J-DAPT 在几乎无额外开销的情况下将检测准确率提升至接近100%。结果表明,J-DAPT 可为 VLM 在机器人应用中的安全性提供实用保障。更多资料见:https://j-dapt.github.io。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly deployed in robotic environments but remain vulnerable to jailbreaking attacks that bypass safety mechanisms and drive unsafe or physically harmful behaviors in the real world. Data-driven defenses such as jailbreak classifiers show promise, yet they struggle to generalize in domains where specialized datasets are scarce, limiting their effectiveness in robotics and other safety-critical contexts. To address this gap, we introduce J-DAPT, a lightweight framework for multimodal jailbreak detection through attention-based fusion and domain adaptation. J-DAPT integrates textual and visual embeddings to capture both semantic intent and environmental grounding, while aligning general-purpose jailbreak datasets with domain-specific reference data. Evaluations across autonomous driving, maritime robotics, and quadruped navigation show that J-DAPT boosts detection accuracy to nearly 100% with minimal overhead. These results demonstrate that J-DAPT provides a practical defense for securing VLMs in robotic applications. Additional materials are made available at: https://j-dapt.github.io.

机器人安全越狱检测多模态域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。