arXiv:2508.21456cs.HCcs.CL2025-08被引 32

让界面代理在关键选择处暂停,让用户掌控决策。

Morae: Proactively Pausing UI Agents for User Choices

  • 通过多模态模型分析用户指令与界面信息,识别需用户决策的节点。
  • 实测中帮助视障用户完成更多任务,选项更符合个人偏好。
  • 适合关注人机协作、无障碍设计的研究者与开发者。

界面代理有望为视障和低视力(BLV)用户简化难以访问或复杂的界面操作。然而,当前的界面代理通常以端到端方式执行任务,未在关键决策环节引入用户参与或提醒重要上下文信息,削弱了用户的自主性。例如,在一项实地研究中,一位视障用户希望购买最便宜的起泡水,代理自动从多个价格相同的选项中选择一个,却未提及不同口味或更高评分的替代品。为此,我们提出 Morae,一种能在任务执行过程中自动识别决策点并暂停的界面代理,以便用户做出选择。Morae 利用大型多模态模型结合用户查询、界面代码和截图,当存在选择时主动向用户提问以获取澄清。在针对真实网络任务的实验中,与基线代理(包括 OpenAI Operator)相比,Morae 显著提升了用户任务完成率,并使所选选项更契合其偏好。本工作展示了混合主动性交互范式:用户既享受界面代理的自动化便利,又能表达自身偏好。

原文摘要 · Abstract (English)

User interface (UI) agents promise to make inaccessible or complex UIs easier to access for blind and low-vision (BLV) users. However, current UI agents typically perform tasks end-to-end without involving users in critical choices or making them aware of important contextual information, thus reducing user agency. For example, in our field study, a BLV participant asked to buy the cheapest available sparkling water, and the agent automatically chose one from several equally priced options, without mentioning alternative products with different flavors or better ratings. To address this problem, we introduce Morae, a UI agent that automatically identifies decision points during task execution and pauses so that users can make choices. Morae uses large multimodal models to interpret user queries alongside UI code and screenshots, and prompt users for clarification when there is a choice to be made. In a study over real-world web tasks with BLV participants, Morae helped users complete more tasks and select options that better matched their preferences, as compared to baseline agents, including OpenAI Operator. More broadly, this work exemplifies a mixed-initiative approach in which users benefit from the automation of UI agents while being able to express their preferences.

界面代理视障辅助人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。