arXiv:2508.19493cs.CRcs.CV2025-08AAAI被引 11

首个大规模测试手机AI助手隐私意识的基准,发现多数表现不佳。

Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents

  • 构建7138个场景的隐私基准,标注信息类型与敏感度。
  • 主流助手隐私识别率普遍低于60%,闭源模型优于开源。
  • 敏感度越高越易被识别,提示也难提升效果。

智能手机为用户带来便利的同时,也广泛记录各类个人信息。当前基于多模态大语言模型(MLLM)的手机助手在任务自动化上表现优异,但其操作中获得了对敏感用户信息的广泛访问权限。为全面评估此类助手的隐私意识,我们首次提出涵盖7,138个场景的大规模基准,对每个场景的隐私类型(如账户凭证)、敏感等级和位置进行标注。我们系统评测了七款主流手机助手。结果表明,几乎所有被测助手的隐私意识(RA)均不理想,即使有明确提示,性能仍低于60%。闭源模型整体优于开源模型,其中Gemini 2.0-flash表现最佳,达67%。此外,助手的隐私检测能力与场景敏感度高度相关:敏感度越高,越容易被识别。研究呼吁学界重新思考手机助手在功能与隐私间的失衡问题。代码与基准数据已公开于https://zhixin-l.github.io/SAPA-Bench。

原文摘要 · Abstract (English)

Smartphones bring significant convenience to users but also enable devices to extensively record various types of personal information. Existing smartphone agents powered by Multimodal Large Language Models (MLLMs) have achieved remarkable performance in automating different tasks. However, as the cost, these agents are granted substantial access to sensitive users' personal information during this operation. To gain a thorough understanding of the privacy awareness of these agents, we present the first large-scale benchmark encompassing 7,138 scenarios to the best of our knowledge. In addition, for privacy context in scenarios, we annotate its type (e.g., Account Credentials), sensitivity level, and location. We then carefully benchmark seven available mainstream smartphone agents. Our results demonstrate that almost all benchmarked agents show unsatisfying privacy awareness (RA), with performance remaining below 60% even with explicit hints. Overall, closed-source agents show better privacy ability than open-source ones, and Gemini 2.0-flash achieves the best, achieving an RA of 67%. We also find that the agents' privacy detection capability is highly related to scenario sensitivity level, i.e., the scenario with a higher sensitivity level is typically more identifiable. We hope the findings enlighten the research community to rethink the unbalanced utility-privacy tradeoff about smartphone agents. Our code and benchmark are available at https://zhixin-l.github.io/SAPA-Bench.

隐私安全大模型手机助手基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。