arXiv:2603.26648cs.SEcs.AI2026-03被引 9

构建首个分层视觉网页开发评测基准,验证代码代理真实能力

Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

  • 从界面到代码的全流程分层评测,覆盖静态生成到全栈开发
  • 含193项任务、918张原型图和1255个测试用例,基于真实网站构建
  • 提出双组件验证机制,可精准评估模型在复杂任务中的表现

大语言模型提升了代码代理的能力,但对复杂端到端网页开发的系统性评估仍显不足。为此,我们提出 Vision2Web,一个分层的视觉网页开发评测基准,涵盖从静态界面转代码、交互式多页面前端复现到长周期全栈开发。该基准基于真实网站,包含16类共193项任务,配有918张原型图像和1255个测试用例。为支持灵活、全面且可靠的评估,我们提出基于工作流的代理验证范式,包含图形界面代理验证器与视觉语言模型判别器两个互补组件。我们在不同编码代理框架下评估多个视觉语言模型,发现各层级均存在显著性能差距,当前最先进模型在全栈开发任务上仍表现不佳。

原文摘要 · Abstract (English)

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical benchmark for visual website development, spanning from static UI-to-code generation, interactive multi-page frontend reproduction, to long-horizon full-stack website development. The benchmark is constructed from real-world websites and comprises a total of 193 tasks across 16 categories, with 918 prototype images and 1,255 test cases. To support flexible, thorough and reliable evaluation, we propose workflow-based agent verification paradigm based on two complementary components: a GUI agent verifier and a VLM-based judge. We evaluate multiple visual language models instantiated under different coding-agent frameworks, revealing substantial performance gaps at all task levels, with state-of-the-art models still struggling on full-stack development.

网页生成评测基准代码代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。