arXiv:2507.02306cs.HCcs.AI2025-07被引 5

用AI替代人工做界面可用性评估,效果更稳且发现更多布局问题。

Synthetic Heuristic Evaluation: A Comparison between AI- and Human-Powered Usability Evaluation

  • 用多模态大模型分析界面图并给出设计反馈
  • 识别出73%~77%的可用性问题,优于5名人类专家
  • 适合需要快速、稳定评估的UI设计团队

可用性评估在以用户为中心的设计中至关重要,但成本高昂,需专家时间和用户补偿。本文提出一种基于多模态大语言模型的合成启发式评估方法,可分析界面图像并提供设计反馈。在两款应用上对比经验丰富的用户体验专家,该方法分别识别出73%和77%的可用性问题,高于5名人类评估者的平均表现(57%和63%)。相比人类评估,合成评估在任务间表现更一致,尤其擅长发现布局问题,显示出潜在的注意力与感知优势。然而,在识别部分UI组件、设计规范及跨屏违规方面仍存在不足。长期测试表明其性能稳定。本研究揭示了人与大模型评估间的差异,为合成启发式评估的设计提供了依据。

原文摘要 · Abstract (English)

Usability evaluation is crucial in human-centered design but can be costly, requiring expert time and user compensation. In this work, we developed a method for synthetic heuristic evaluation using multimodal LLMs' ability to analyze images and provide design feedback. Comparing our synthetic evaluations to those by experienced UX practitioners across two apps, we found our evaluation identified 73% and 77% of usability issues, which exceeded the performance of 5 experienced human evaluators (57% and 63%). Compared to human evaluators, the synthetic evaluation's performance maintained consistent performance across tasks and excelled in detecting layout issues, highlighting potential attentional and perceptual strengths of synthetic evaluation. However, synthetic evaluation struggled with recognizing some UI components and design conventions, as well as identifying across screen violations. Additionally, testing synthetic evaluations over time and accounts revealed stable performance. Overall, our work highlights the performance differences between human and LLM-driven evaluations, informing the design of synthetic heuristic evaluations.

AI评估可用性测试多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。