arXiv:2505.05026cs.CLcs.LG2025-05ACL被引 2

构建首个评估MLLM理解界面如何影响用户行为的基准

Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding

  • 基于300组真实A/B测试界面,构建可量化行为影响的评估基准
  • 多模型测试显示现有MLLM对设计影响用户行为的理解有限
  • 提供专家解读支持,适合研究人机交互与视觉设计的学者

用户界面(UI)设计不仅关乎视觉呈现,更深刻影响用户体验(UX),推动了UI/UX一体化的发展。尽管近期研究利用多模态大语言模型(MLLMs)评估界面,但大多仅关注表面特征,忽视设计选择对用户行为的规模化影响。为此,我们提出WiserUI-Bench,一个全新的基准,用于多模态理解界面与用户体验如何影响用户行为。该基准基于300组来自工业界A/B测试的真实界面图像对,其胜出方案经实证验证能引发更多用户操作。为支持实践中的设计改进,我们还提供专家精心标注的关键解释。在多个MLLM上针对两项任务进行实验:(1) 预测一对A/B测试界面中哪个更有效;(2) 生成与专家解释对齐的后验解释。结果表明,当前模型对界面设计影响用户行为的理解仍十分有限。我们认为本工作将推动利用MLLM在用户行为语境下进行视觉设计研究。

原文摘要 · Abstract (English)

User interface (UI) design goes beyond visuals to shape user experience (UX), underscoring the shift toward UI/UX as a unified concept. While recent studies have explored UI evaluation using Multimodal Large Language Models (MLLMs), they largely focus on surface-level features, overlooking how design choices influence user behavior at scale. To fill this gap, we introduce WiserUI-Bench, a novel benchmark for multimodal understanding of how UI/UX design affects user behavior, built on 300 real-world UI image pairs from industry A/B tests, with empirically validated winners that induced more user actions. For future design progress in practice, post-hoc understanding of why such winners succeed with mass users is also required; we support this via expert-curated key interpretations for each instance. Experiments across multiple MLLMs on WiserUI-Bench for two main tasks, (1) predicting the more effective UI image between an A/B-tested pair, and (2) explaining it post-hoc in alignment with expert interpretations, show that models exhibit limited understanding of the behavioral impact of UI/UX design. We believe our work will foster research on leveraging MLLMs for visual design in user behavior contexts.

多模态模型界面设计用户行为评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。