arXiv:2603.04601cs.SEcs.AI2026-03中稿 · ACM CAIS 2026被引 7

测试AI从零构建完整网页应用的能力,发现当前最佳模型仅61.8%准确率。

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

  • 用真实浏览器自动化流程评估模型端到端开发能力
  • 16个前沿模型中最高准确率61.8%,自检机制预测性能相关性达0.72
  • 揭示评测者选择显著影响结果,适合研究代码生成与评估方法的团队

代码生成已成为人工智能最具影响力的应用之一,但现有基准多衡量孤立任务,而非从零开始构建可运行应用的完整过程。我们提出Vibe Code Bench,包含100个网页应用需求(50个私有验证集,50个保留测试集),涵盖964个基于浏览器的工作流,共10,131个子步骤,由自主浏览器代理对部署应用进行评估。在16个前沿模型中,表现最好的模型在测试集上达到61.8%的准确率,表明可靠的端到端应用开发仍是重大挑战。我们发现生成过程中的自我测试是性能的重要预测因子(皮尔逊相关系数r=0.72),并通过一项完成的人类对齐研究显示,评估者选择会显著影响结果(成对步骤级一致率在31.8%-93.6%之间)。我们的贡献包括:(1) 一个新颖的基准数据集和基于浏览器的端到端网页应用开发评估流水线;(2) 对16个前沿模型的全面评估,含成本、延迟和错误分析;(3) 一种包含跨模型与人类标注结果的评估者对齐协议。

原文摘要 · Abstract (English)

Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete "zero-to-one" process of building a working application from scratch. We introduce Vibe Code Bench, a benchmark of 100 web application specifications (50 private validation, 50 held-out test) with 964 browser-based workflows comprising 10,131 substeps, evaluated against deployed applications by an autonomous browser agent. Across 16 frontier models, the best achieves 61.8% accuracy on the test split, revealing that reliable end-to-end application development remains a frontier challenge. We identify self-testing during generation as a strong performance predictor (Pearson r=0.72), and show through a completed human alignment study that evaluator selection materially affects outcomes (31.8-93.6% pairwise step-level agreement). Our contributions include (1) a novel benchmark dataset and browser-based evaluation pipeline for end-to-end web application development, (2) a comprehensive evaluation of 16 frontier models with cost, latency, and error analysis, and (3) an evaluator alignment protocol with both cross-model and human annotation results.

代码生成端到端评估智能开发基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。