arXiv:2605.07630cs.CLcs.AI2026-05被引 1

区分手机助手的‘安全’是明智选择还是根本不会操作。

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

论文配图:Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
图 1 · 摘自论文原文
  • 设计新基准PhoneSafety,分离安全决策与操作能力
  • 8个模型测试显示强能力不等于更安全,失败多在复杂界面
  • 两类错误:错选危险操作或根本无法响应,需分别修复

当手机使用代理避免了伤害,这究竟是出于安全意识,还是根本无法操作?现有评估难以区分。有害结果可能因代理识别风险而选择安全动作,也可能因其无法理解屏幕或执行任何操作。这两种情况成因不同,需不同修正,但当前基准常将它们混为任务成功、拒绝或最终有害结果。我们提出PhoneSafety,一个包含700个真实手机交互中关键安全时刻的基准,覆盖130多个应用。每个实例聚焦于高风险时刻的下一步决策,仅问:代理是否采取安全操作、采取危险操作,或完全无有效行为?我们在该框架下评估8个代表性手机使用代理。结果显示两大模式:第一,更强的通用手机操作能力并不意味着高风险时刻更安全;在常规任务表现好的模型,在关键决策中未必更安全。第二,无法有效行动的表现更像是能力不足信号而非安全信号:这些失败集中在视觉和操作更复杂的场景,且在评估协议改变时保持稳定。各模型的失败可归纳为两种重复模式:能操作却误选危险动作,以及在复杂界面中完全无法响应。总体而言,无害结果不足以证明安全。评估手机使用代理必须分离不当判断与操作无能。

原文摘要 · Abstract (English)

When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may be avoided because the agent recognized the risk and chose the safe action, or because it failed to understand the screen or execute any relevant action at all. These cases have different causes and call for different fixes, yet current benchmarks often merge them under task success, refusal, or final harmful outcome. We address this problem with PhoneSafety, a benchmark of 700 safety-critical moments drawn from real phone interactions across more than 130 apps. Each instance isolates the next decision at a risky moment and asks a simple question: does the model take the safe action, take the unsafe action, or fail to do anything useful? We evaluate eight representative phone-use agents under this framework. Our results reveal two main patterns. First, stronger general phone-use ability does not reliably imply safer choices at risky moments. Models that perform better on ordinary app tasks are not always the ones that behave more safely when the next action matters. Second, failures to do anything useful behave like a capability signal rather than a safety signal: they are concentrated in more visually and operationally demanding settings and remain stable when the evaluation protocol changes. Across models, failures split into two recurring patterns: unsafe choices in settings where the model can act but chooses wrongly, and inability to act in more visually and operationally demanding screens. Overall, a harmless outcome is not enough to count as evidence of safety. Evaluating phone-use agents requires separating unsafe judgment from inability to act.

手机代理安全评估行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。