arXiv:2501.17799cs.HCcs.IR2025-01被引 27

用多模态大模型提升移动UI设计灵感搜索的语义理解与相关性。

Leveraging Multimodal LLM for Inspirational User Interface Search

  • 利用多模态大模型从UI图像中提取用户群体、应用情绪等关键语义。
  • 相比现有方法,检索准确率显著提升,人机评估表现更优。
  • 适合需要深度理解设计语境的UI设计师和人机交互研究者。

灵感搜索是移动用户界面(UI)设计中探索参考以激发新创意的关键过程。然而,面对海量的UI参考,如何有效探索仍是一大挑战。现有基于AI的UI搜索方法常忽略目标用户、应用氛围等重要语义信息,且通常依赖视图层级等元数据,限制了实际应用。本文通过一项形成性研究识别出关键的UI语义,并利用多模态大语言模型(MLLM)从移动UI图像中提取与解释这些语义,构建了一个基于语义的UI搜索系统。经计算与人工评估,该方法在多个指标上显著优于现有方法,为UI设计师提供了更丰富、更具上下文相关性的搜索体验。本研究深化了对移动UI设计语义的理解,揭示了MLLM在灵感搜索中的潜力,并提供了可用于未来研究的丰富UI语义数据集。

原文摘要 · Abstract (English)

Inspirational search, the process of exploring designs to inform and inspire new creative work, is pivotal in mobile user interface (UI) design. However, exploring the vast space of UI references remains a challenge. Existing AI-based UI search methods often miss crucial semantics like target users or the mood of apps. Additionally, these models typically require metadata like view hierarchies, limiting their practical use. We used a multimodal large language model (MLLM) to extract and interpret semantics from mobile UI images. We identified key UI semantics through a formative study and developed a semantic-based UI search system. Through computational and human evaluations, we demonstrate that our approach significantly outperforms existing UI retrieval methods, offering UI designers a more enriched and contextually relevant search experience. We enhance the understanding of mobile UI design semantics and highlight MLLMs' potential in inspirational search, providing a rich dataset of UI semantics for future studies.

UI设计多模态模型灵感搜索语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。