arXiv:2506.13832cs.SEcs.AI2025-06被引 13

构建首个自动化评估前端代码生成的基准测试,真实还原开发场景。

FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation

  • 人机协作设计148个从基础到复杂的前端任务
  • 自动评测框架与专家评估达90.54%一致率
  • 适合研究前端生成模型性能及优化方向

大型语言模型在前端代码生成方面取得显著进展,但现有基准存在任务过于简单、测试用例不严谨、缺乏端到端验证等问题,影响模型性能的准确评估。为此,我们提出FrontendBench,一个由人类与大模型共同开发的基准测试。该基准按代码功能分类,融入交互式测试场景,覆盖五个层级的网页组件,从基础界面元素到复杂交互功能,每项任务均反映真实开发挑战。共包含148组精心设计的提示-测试用例对。此外,我们构建了自动评估框架,在沙盒环境中执行生成代码,并通过预设测试脚本评估结果,其与人工专家评估达成90.54%的一致率,证明了可靠性。我们在FrontendBench上对多个先进LLM进行测评,发现其在处理真实前端任务时表现差异显著。结果表明,FrontendBench是可靠且可扩展的基准,支持一致的多模态评估,为未来前端代码生成研究提供坚实基础。数据与代码将尽快公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made significant strides in front-end code generation. However, existing benchmarks exhibit several critical limitations: many tasks are overly simplistic, test cases often lack rigor, and end-to-end validation is absent. These issues hinder the accurate assessment of model performance. To address these challenges, we present FrontendBench, a benchmark co-developed by humans and LLMs. FrontendBench categorizes tasks based on code functionality and incorporates interactive test scenarios, enabling a more comprehensive and practical evaluation of front-end code generation capabilities. The benchmark comprises 148 meticulously crafted prompt-test case pairs spanning five levels of web components, from basic UI elements to complex interactive features. Each task reflects realistic front-end development challenges. Furthermore, we introduce an automatic evaluation framework that executes generated code within a sandbox environment and assesses outcomes using predefined test scripts. This framework achieves a 90.54% agreement rate with expert human evaluations, demonstrating high reliability. We benchmark several state-of-the-art LLMs on FrontendBench and observe substantial performance disparities in handling real-world front-end tasks. These results highlight FrontendBench as a reliable and scalable benchmark, supporting consistent multimodal evaluation and providing a robust foundation for future research in front-end code generation. Our data and code will be released soon.

前端生成自动评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。