arXiv:2604.26020cs.CLcs.AI2026-04被引 2

用AI代理自动评估界面易用性,比大模型更准更像人。

Training Computer Use Agents to Assess the Usability of Graphical User Interfaces

论文配图:Training Computer Use Agents to Assess the Usability of Graphical User Interfaces
图 1 · 摘自论文原文
  • 基于重要操作流程训练代理,模拟人类交互行为。
  • 在真实和合成界面中评分准确率超越更大模型。
  • 适合人机交互研究者与产品设计团队快速验证界面。

使用专家和潜在用户进行可用性测试可评估图形用户界面(GUI)的有效性、效率和用户满意度,但该过程成本高且耗时。以往研究采用计算机使用代理(CUAs)等生成式代理来模拟用户交互与偏好,但我们发现这些代理仍难以提供准确的可用性评估。本文提出一种新型机器学习方法,将可用性概念转化为可计算形式,通过三步训练计算机使用代理:一、优先处理关键交互路径;二、以类人方式执行操作;三、预测学习到的数值化可用性得分。我们在大规模带可用性标签和人类偏好的交互界面数据集上训练了名为 uxCUA 的代理。结果表明,uxCUA 在准确评估可用性方面优于更大模型,并能对合成与真实界面生成合理批评。本工作旨在为人机交互领域的自动化可用性评估建立严谨的数据驱动基础。

原文摘要 · Abstract (English)

Usability testing with experts and potential users can assess the effectiveness, efficiency, and user satisfaction of graphical user interfaces (GUIs) but doing so remains a costly and time-intensive process. Prior work has used computer use agents (CUAs) and other generative agents that can simulate user interactions and preference, but we show that agents still struggle to provide accurate usability assessments. In this work, we present a novel machine learning method that operationalizes a computational definition of usability to train CUAs to assess GUI usability by i) prioritizing important interaction flows, ii) executing them through human-like interactions, and iii) predicting a learned numerical usability score. We train a computer use agent, uxCUA, with our algorithm on a large-scale dataset of fully interactive user interfaces (UIs) paired with usability labels and human preferences. We show that uxCUA outperforms larger models in accurate usability assessments and produces realistic critiques of both synthetic and real UIs. More broadly, our work aims to build a principled, data-driven foundation for automated usability assessment in HCI.

可用性评估智能代理人机交互UI测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。