为可信的GUI智能体设计以人为本的评估框架,解决隐私安全短板
Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents
- 提出融合风险评估与用户知情同意的评估思路
- 指出现有评测忽视隐私安全,仅关注性能表现
- 适合关注AI代理安全性的研究者与开发者
大型语言模型(LLMs)推动了基于LLM的图形界面(GUI)自动化发展,但其在缺乏充分人类监管下处理敏感数据,带来显著的隐私与安全风险。本文指出GUI智能体存在三大关键风险,并揭示其与传统GUI自动化及通用自主代理的差异。尽管如此,现有评估仍主要聚焦性能,对隐私与安全的考量严重不足。我们综述了当前针对GUI及通用LLM智能体的评估指标,归纳出五项将人类评估者纳入GUI智能体测评的核心挑战。为弥补这些空白,我们倡导构建以人为本的评估框架,整合风险评估、通过上下文提示增强用户知情权,并将隐私与安全嵌入智能体的设计与评测流程中。
原文摘要 · Abstract (English)
The rise of Large Language Models (LLMs) has revolutionized Graphical User Interface (GUI) automation through LLM-powered GUI agents, yet their ability to process sensitive data with limited human oversight raises significant privacy and security risks. This position paper identifies three key risks of GUI agents and examines how they differ from traditional GUI automation and general autonomous agents. Despite these risks, existing evaluations focus primarily on performance, leaving privacy and security assessments largely unexplored. We review current evaluation metrics for both GUI and general LLM agents and outline five key challenges in integrating human evaluators for GUI agent assessments. To address these gaps, we advocate for a human-centered evaluation framework that incorporates risk assessments, enhances user awareness through in-context consent, and embeds privacy and security considerations into GUI agent design and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。