arXiv:2506.07818cs.CL2025-06ACL被引 16

构建多维度网页生成评估基准,揭示大模型在代码开发中的真实短板。

WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code

  • 从软件工程出发,分四个维度系统评测大模型网页开发能力。
  • 基于2.1万条真实网页数据,覆盖0.7千个真实网站的高质量问答对。
  • 首次揭示模型在界面理解与代码生成间的断层问题,适合研究开发者参考。

随着生成式AI技术的快速发展,多模态大语言模型(MLLMs)有望成为能够执行复杂网页应用开发的AI工程师。由于模型需融合多维度子能力应对不同开发阶段挑战,构建多视角评估框架对准确引导开发效率提升至关重要。然而,现有基准通常仅关注网页生成结果,缺乏对子能力的细致评估。本文借鉴软件工程原则,提出WebUIBench,一个系统化设计的基准,用于评估MLLMs在四个关键领域的能力:网页界面感知、HTML编程、界面-代码理解以及网页到代码转换。该基准包含21,000条高质量问答对,源自超过700个真实网站。对29个主流MLLMs的广泛评估揭示了模型在开发过程中表现出的技能特征与各类缺陷。

原文摘要 · Abstract (English)

With the rapid advancement of Generative AI technology, Multimodal Large Language Models(MLLMs) have the potential to act as AI software engineers capable of executing complex web application development. Considering that the model requires a confluence of multidimensional sub-capabilities to address the challenges of various development phases, constructing a multi-view evaluation framework is crucial for accurately guiding the enhancement of development efficiency. However, existing benchmarks usually fail to provide an assessment of sub-capabilities and focus solely on webpage generation outcomes. In this work, we draw inspiration from the principles of software engineering and further propose WebUIBench, a benchmark systematically designed to evaluate MLLMs in four key areas: WebUI Perception, HTML Programming,WebUI-HTML Understanding, and WebUI-to-Code. WebUIBench comprises 21K high-quality question-answer pairs derived from over 0.7K real-world websites. The extensive evaluation of 29 mainstream MLLMs uncovers the skill characteristics and various weakness that models encountered during the development process.

多模态模型代码生成评估基准网页开发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。