arXiv:2410.17520cs.LGcs.CL2024-10AAAI被引 48

首个评估手机控制智能体安全性的基准,发现主流模型易出错。

MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control

  • 用安卓模拟器构建真实手机环境测试智能体行为
  • 基线模型在银行、消息等任务中频繁引发安全隐患
  • 提出安全优先提示法,但仍需改进以获用户信任

由大语言模型驱动的自主智能体在辅助任务中展现出巨大潜力,尤其在手机设备控制领域。由于这些智能体直接接触个人数据和系统设置,确保其安全可靠至关重要,以防不良后果。然而,目前尚无针对手机控制智能体安全性的标准化评估基准。本文提出 MobileSafetyBench,一个基于安卓模拟器的基准平台,用于评估智能体在真实移动环境中的安全性。我们设计了涵盖消息、银行等应用的多样化任务,测试智能体在日常使用场景中的表现及其对间接提示注入攻击的鲁棒性。实验表明,基于先进大模型的基线智能体在执行任务时常无法有效防止危害。为此,我们提出一种鼓励智能体优先考虑安全性的提示方法,虽有一定成效,但距离建立用户完全信任仍有较大差距。这凸显了在移动环境中持续研发更稳健安全机制的紧迫性。

原文摘要 · Abstract (English)

Autonomous agents powered by large language models (LLMs) show promising potential in assistive tasks across various domains, including mobile device control. As these agents interact directly with personal information and device settings, ensuring their safe and reliable behavior is crucial to prevent undesirable outcomes. However, no benchmark exists for standardized evaluation of the safety of mobile device-control agents. In this work, we introduce MobileSafetyBench, a benchmark designed to evaluate the safety of device-control agents within a realistic mobile environment based on Android emulators. We develop a diverse set of tasks involving interactions with various mobile applications, including messaging and banking applications, challenging agents with managing risks encompassing misuse and negative side effects. These tasks include tests to evaluate the safety of agents in daily scenarios as well as their robustness against indirect prompt injection attacks. Our experiments demonstrate that baseline agents, based on state-of-the-art LLMs, often fail to effectively prevent harm while performing the tasks. To mitigate these safety concerns, we propose a prompting method that encourages agents to prioritize safety considerations. While this method shows promise in promoting safer behaviors, there is still considerable room for improvement to fully earn user trust. This highlights the urgent need for continued research to develop more robust safety mechanisms in mobile environments.

智能体安全LLM评估移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。