arXiv:2604.20279cs.HCcs.AI2026-04

AgentLens让手机智能助手根据任务自动切换显示方式,提升人机协作体验。

AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents

论文配图:AgentLens: Adaptive Visual Modalities for Human-Agent Interaction in Mobile GUI Agents
图 1 · 摘自论文原文
  • 根据任务动态选择全界面、局部界面或生成界面三种显示模式。
  • 在21人实验中用户满意度达85.7%,可用性评分1.94,意向分6.43/7。
  • 支持后台运行并仅关键时显示可视化,兼顾透明与多任务效率。

移动GUI代理可通过直接操作应用界面自动化手机任务,但其执行过程中如何与用户沟通仍缺乏研究。现有系统依赖两种极端:前景执行虽透明度高但妨碍多任务,背景执行虽支持多任务却缺乏视觉感知。通过迭代式前期研究,我们发现用户更倾向即时响应的混合模式,而最佳可视化方式取决于任务类型。为此,我们提出AgentLens,一种可自适应使用三种视觉模态(全界面、局部界面、GenUI)的移动GUI代理。AgentLens在标准移动代理基础上增加自适应通信动作,并利用虚拟显示实现后台运行与选择性视觉叠加。在21名参与者控制实验中,85.7%用户偏好AgentLens,整体可用性得分为1.94(PSSUQ),采纳意愿达6.43/7。

原文摘要 · Abstract (English)

Mobile GUI agents can automate smartphone tasks by interacting directly with app interfaces, but how they should communicate with users during execution remains underexplored. Existing systems rely on two extremes: foreground execution, which maximizes transparency but prevents multitasking, and background execution, which supports multitasking but provides little visual awareness. Through iterative formative studies, we found that users prefer a hybrid model with just-in-time visual interaction, but the most effective visualization modality depends on the task. Motivated by this, we present AgentLens, a mobile GUI agent that adaptively uses three visual modalities during human-agent interaction: Full UI, Partial UI, and GenUI. AgentLens extends a standard mobile agent with adaptive communication actions and uses Virtual Display to enable background execution with selective visual overlays. In a controlled study with 21 participants, AgentLens was preferred by 85.7% of participants and achieved the highest usability (1.94 Overall PSSUQ) and adoption-intent (6.43/7).

人机交互移动代理自适应界面可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。