arXiv:2510.08783cs.HCcs.AI2025-10被引 13

用大模型评估界面设计,看它能否像人一样判断好坏。

MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces

  • 用众包数据对比大模型与人类对30个界面的评分
  • 部分维度如美观度预测准确,但交互体验差异明显
  • 适合早期设计筛选,不适合替代真实用户测试

理想的设计流程中,用户界面(UI)设计应与用户研究结合以验证决策,但在早期探索阶段常受资源限制。近年来多模态大语言模型(MLLMs)的发展为早期评估提供了可能,帮助设计师在正式测试前缩小选择范围。不同于以往聚焦电商等特定领域点击率或转化率等行为指标的研究,本文关注跨场景界面的主观用户评价。我们探究了MLLMs在评估单个界面及比较界面时,是否能模拟人类偏好。基于众包平台数据,我们在30个界面中对GPT-4o、Claude和Llama进行了基准测试,分析其与人类判断在多个UI维度上的对齐程度。结果显示,某些维度上MLLMs可近似人类偏好,但在其他维度存在显著偏差,凸显其在辅助早期用户体验研究中的潜力与局限。

原文摘要 · Abstract (English)

In an ideal design pipeline, user interface (UI) design is intertwined with user research to validate decisions, yet studies are often resource-constrained during early exploration. Recent advances in multimodal large language models (MLLMs) offer a promising opportunity to act as early evaluators, helping designers narrow options before formal testing. Unlike prior work that emphasizes user behavior in narrow domains such as e-commerce with metrics like clicks or conversions, we focus on subjective user evaluations across varied interfaces. We investigate whether MLLMs can mimic human preferences when evaluating individual UIs and comparing them. Using data from a crowdsourcing platform, we benchmark GPT-4o, Claude, and Llama across 30 interfaces and examine alignment with human judgments on multiple UI factors. Our results show that MLLMs approximate human preferences on some dimensions but diverge on others, underscoring both their potential and limitations in supplementing early UX research.

界面评估多模态模型用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。