arXiv:2507.00841cs.AIcs.CR2025-07被引 13

提出链式越狱检测与自动化评估框架,提升多模态移动智能体安全性。

SafeMobile: Chain-level Jailbreak Detection and Automated Evaluation for Multimodal Mobile Agents

  • 基于行为序列构建链级风险识别机制
  • 在高危任务中显著提升越狱行为检出率
  • 适合安全评测与智能代理防护研究者参考

随着多模态基础模型在智能代理系统中的广泛应用,手机设备控制、智能助手交互及多模态任务执行等场景正逐步依赖大模型驱动的代理系统。然而,这些系统也面临日益突出的越狱风险:攻击者可通过特定输入诱导代理绕过原始行为约束,进而触发修改设置、执行未授权命令或冒充用户身份等敏感操作,带来新的安全挑战。现有智能代理安全措施在复杂交互场景下仍存在局限,尤其难以有效检测多轮对话或任务序列中的潜在风险行为。此外,缺乏高效且一致的自动化评估方法来辅助判断此类风险的影响。本文探讨了多模态移动智能代理的安全问题,提出结合行为序列信息的风险判别机制,并设计基于大语言模型的自动化辅助评估方案。在多个典型高危任务中的初步验证表明,该方法可一定程度上提升对风险行为的识别能力,有助于降低代理被越狱的概率。本研究旨在为多模态智能代理系统的安全风险建模与防护提供有益参考。

原文摘要 · Abstract (English)

With the wide application of multimodal foundation models in intelligent agent systems, scenarios such as mobile device control, intelligent assistant interaction, and multimodal task execution are gradually relying on such large model-driven agents. However, the related systems are also increasingly exposed to potential jailbreak risks. Attackers may induce the agents to bypass the original behavioral constraints through specific inputs, and then trigger certain risky and sensitive operations, such as modifying settings, executing unauthorized commands, or impersonating user identities, which brings new challenges to system security. Existing security measures for intelligent agents still have limitations when facing complex interactions, especially in detecting potentially risky behaviors across multiple rounds of conversations or sequences of tasks. In addition, an efficient and consistent automated methodology to assist in assessing and determining the impact of such risks is currently lacking. This work explores the security issues surrounding mobile multimodal agents, attempts to construct a risk discrimination mechanism by incorporating behavioral sequence information, and designs an automated assisted assessment scheme based on a large language model. Through preliminary validation in several representative high-risk tasks, the results show that the method can improve the recognition of risky behaviors to some extent and assist in reducing the probability of agents being jailbroken. We hope that this study can provide some valuable references for the security risk modeling and protection of multimodal intelligent agent systems.

智能代理越狱检测多模态安全自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。