评测模型理解多语言网页并生成代码的能力,发现现有模型在推理和代码保持功能上仍有不足。
WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- 统一视觉问答、代码编辑与原型转代码三任务,用真实网页数据评估
- 模型在复杂推理和代码功能保持上表现差,尤其多语言内容处理弱
- 适合研究多模态智能体与跨语言网页自动化开发的学者参考
我们提出WebMMU,一个面向多语言网页理解与代码生成的基准测试。该基准涵盖三大核心任务:(1)网站视觉问答,(2)涉及HTML/CSS/JavaScript的代码编辑,(3)原型到代码生成。与以往将任务割裂的基准不同,WebMMU采用专家标注的真实网页数据,统一评估模型在多步推理、精确元素定位、功能化界面理解与编码方面的能力。评估显示,尽管多模态大模型在基础信息提取上表现良好,但在推理、元素定位、保持代码功能完整性以及生成具有层级结构并支持多语言内容的代码方面仍存在明显短板。这些结果揭示了当前多模态大模型的关键局限,凸显了提升多模态与跨语言推理能力的必要性,以构建能自动化完成多样化网页开发任务的未来网络智能体。
原文摘要 · Abstract (English)
We present WebMMU, a multilingual benchmark that evaluates three core web tasks: (1) website visual question answering, (2) code editing involving HTML/CSS/JavaScript, and (3) mockup-to-code generation. Unlike prior benchmarks that treat these tasks separately, WebMMU unifies them using expert-annotated, real-world web data to assess models' abilities in complex multi-step reasoning, precise element grounding, and functional UI comprehension and coding. Our evaluation shows that while multimodal large language models (MLLMs) perform well on basic information extraction, they struggle with reasoning and grounding, editing code to preserve functionality, and generating design-to-code that maintains hierarchy and supports multilingual content. These findings reveal key limitations in current MLLMs and underscore the need for improved multimodal and cross-lingual reasoning to build future web agents capable of automating diverse web development tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。