arXiv:2508.16165cs.SEcs.AI2025-08被引 1

用多模态大模型辅助评估界面可用性,提升效率与可及性。

Investigating Multimodal Large Language Models to Support Usability Evaluation

  • 将可用性评估任务转化为问题优先级排序,结合图文信息分析。
  • 多模型生成结果与专家评估高度一致,能有效识别关键问题。
  • 提供可视化工具透明化模型输出,适合开发团队快速集成使用。

可用性评估是设计高效直观用户界面的重要方法,但传统方式依赖资源密集型的专家评审,限制了小组织的应用。近期多模态大语言模型(MLLMs)可通过分析文本指令与视觉界面上下文,支持可用性评估。本文将该任务定义为优先级排序问题,识别并解释可用性缺陷,按严重程度排序。研究对比了多个MLLM生成的评估结果与专家判断,结果表明MLLM能提供互补见解,有效支持关键问题的优先处理。此外,本文提出一个交互式可视化工具,便于审查和验证模型输出。基于此,我们构建了将MLLM驱动的可用性评估融入实际开发流程的概念框架。

原文摘要 · Abstract (English)

Usability evaluation is an essential method to support the design of effective and intuitive user interfaces (UIs). However, it commonly relies on resource-intensive, expert-driven methods, which limit its accessibility, especially for small organizations. Recent multimodal large language models (MLLMs) have the potential to support usability evaluation by analyzing textual instructions together with visual UI context. This paper investigates the use of MLLMs as assistive tools for usability evaluation by framing the task as a prioritization problem. It identifies and explains usability issues and ranks them by severity. We report a study that compares the evaluations generated by multiple MLLMs with assessments from usability experts. The results demonstrate that MLLMs can offer complementary insights and support the efficient prioritization of critical issues. Additionally, we present an interactive visualization tool that enables the transparent review and validation of model-generated findings. Based on this, we outline concepts for integrating MLLM-based usability evaluation into real-world development workflows.

多模态模型可用性评估人机交互AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。