测试安卓界面智能体在真实环境攻击下的安全漏洞。
MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps

- 构建142个真实应用任务,用规则+大模型双阶段验证安全
- 六款智能体攻击成功率40.4%至66.9%,普遍不抗干扰
- 适合研究移动AI安全、对抗样本防御的学者和开发者
基于大模型的安卓界面智能体正从研究走向实际应用,但因需处理不可信的环境内容,极易受环境注入攻击影响,包括间接提示注入和恶意指令。此类攻击通过日常手机使用场景悄然改变智能体行为,且现有评估基准多忽略真实用户场景,缺乏对移动端界面智能体在环境注入攻击下的系统性评测。为此,我们提出MobileWorldSafety,一个基于真实安卓应用的142项风险任务基准。每项任务定义可程序验证的风险指标,采用两阶段评估流程:规则引擎处理明确情形,大模型裁判处理模糊情形。该方法区分安全失效与能力失效,实现客观可复现评估。对六种智能体(含通用与专用界面智能体)的测试表明,所有智能体均高度脆弱,攻击成功率达40.4%至66.9%。结果说明当前智能体在遭遇伪装成正常手机上下文的对抗内容时,难以保持安全对齐。MobileWorldSafety为量化此类漏洞提供了基础,推动更鲁棒的移动界面智能体研究。
原文摘要 · Abstract (English)
LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deployment. However, because these agents routinely process untrusted environmental content, they are highly vulnerable to environmental injection attacks, which include indirect prompt injections and adversarial instructions. Such attacks can manipulate the behavior of agents without user awareness through diverse channels encountered in everyday mobile use. Despite these risks, existing benchmarks often fail to capture everyday user scenarios, lacking a systematic evaluation of GUI agents under environmental injection attacks on mobile devices. To address this gap, we introduce MobileWorldSafety, a benchmark of 142 risk tasks built on real Android applications. For each task, we define a programmatically verifiable risk indicator over the final system state and evaluate outcomes with a two-stage pipeline: rule-based verification handles unambiguous cases, while an LLM judge adjudicates ambiguous ones. This distinguishes safety failures from capability failures and enables objective and reproducible assessment. Evaluations on six agents, including both general agents and specialized GUI agents, demonstrate that all agents remain highly vulnerable, with attack success rates ranging from 40.4% to 66.9%. These findings indicate that current agents often fail to maintain safety alignment when adversarial content is presented as ordinary mobile context. MobileWorldSafety provides a foundation for quantifying these vulnerabilities and advancing research on robust mobile GUI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。