arXiv:2603.25864cs.CVcs.AI2026-03中稿 · CVPR被引 3

构建首个评估GUI任务中用户意图的基准,推动人机协作而非单纯自动化。

GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI Tasks

  • 通过屏幕录像与实时口述,收集120名新手在10款软件中的操作数据
  • 模型在行为识别和帮助预测上准确率仅44.6%和55.0%,表现不佳
  • 加入用户上下文可使帮助预测准确率提升50.2个百分点,凸显理解意图的重要性

图形用户界面(GUI)智能体有望协助用户操作复杂软件(如PowerPoint、Photoshop)。现有研究多聚焦于通过点击和按键实现自动化,却忽视了用户的意图表达——用户更希望探索、迭代并自主掌控创作过程。为从自动化转向协作,智能体需理解用户在做什么以及为什么这么做。本文提出GUIDE(GUI用户意图检测评估基准),用于评测模型在开放性GUI任务中感知行为、推断意图并提供适时协助的能力。GUIDE包含120名新手在10种软件上的67.5小时屏幕录像及思考口述数据。该基准定义三项任务:(i) 行为状态检测,(ii) 意图预测,(iii) 帮助预测,分别测试模型对当前行为、目标推理和干预时机判断的能力。在8个先进多模态模型上的评估显示,其行为状态和帮助预测准确率分别仅为44.6%和55.0%。但引入用户上下文后,帮助预测准确率最高提升50.2个百分点,表明结构化理解用户是有效辅助的关键。数据集已公开:https://guide-bench.github.io。

原文摘要 · Abstract (English)

Graphical User Interface (GUI) agents have the potential to assist users in interacting with complex software (e.g., PowerPoint, Photoshop). While prior research has primarily focused on automating user actions through clicks and keystrokes, this paradigm overlooks human intention, where users value the ability to explore, iterate, and refine their ideas while maintaining agency. To move beyond automation and toward collaboration, GUI agents must understand what users are doing and why. We introduce GUIDE (GUI User Intent Detection Evaluation), a benchmark that evaluates AI models on their ability to perceive user behavior, infer intent, and provide assistance in open-ended GUI tasks. GUIDE consists of 67.5 hours of screen recordings from 120 novice user demonstrations with think-aloud narrations, across 10 software. GUIDE defines three tasks - (i) Behavior State Detection, (ii) Intent Prediction, and (iii) Help Prediction that test a model's ability to recognize behavior state, reason about goals, and decide when and how to help. Evaluations across eight state-of-the-art multimodal models reveal that all models struggled, achieving only 44.6% and 55.0% accuracy on behavior state and help prediction. However, providing user context significantly improved the performance, raising help prediction by up to 50.2pp, highlighting the critical role of structured user understanding in effective assistance. Our dataset is available at https://guide-bench.github.io.

GUI智能体用户意图人机协作多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。