首个真实用户需求的网页生成评估基准,支持自动、可解释的多维度评测。
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
- 基于1572条真实用户需求,构建覆盖多模态和表达风格的评估集。
- 24项细粒度指标从9个维度自动评估,结合规则与LLM判别实现客观打分。
- 采用人类偏好加权,结果可解释,适合模型开发者针对性优化。
网页应用已成为大语言模型(LLMs)展示代码生成能力与商业潜力的重要领域。然而,构建面向LLM生成网页应用的基准仍面临挑战:需真实用户需求、无需依赖真实实现或测试用例的通用评估指标,以及可解释的评估结果。为此,我们提出WebCoderBench,首个基于真实采集、可泛化且可解释的网页应用生成评估基准。该基准包含1,572条真实用户需求,涵盖多样化的模态与表达风格,反映真实用户意图。提供24项细粒度评估指标,覆盖9个评估维度,融合规则引擎与LLM作为裁判的范式,实现全自动化、客观且通用的评估。此外,采用人类偏好对齐的权重分配,生成可解释的整体评分。在12个代表性大语言模型与2个基于大模型的智能体上的实验表明,无任一模型在所有指标上占优,为模型开发者提供针对性优化机会。
原文摘要 · Abstract (English)
Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential. However, building a benchmark for LLM-generated web apps remains challenging due to the need for real-world user requirements, generalizable evaluation metrics without relying on ground-truth implementations or test cases, and interpretable evaluation results. To address these challenges, we introduce WebCoderBench, the first real-world-collected, generalizable, and interpretable benchmark for web app generation. WebCoderBench comprises 1,572 real user requirements, covering diverse modalities and expression styles that reflect realistic user intentions. WebCoderBench provides 24 fine-grained evaluation metrics across 9 perspectives, combining rule-based and LLM-as-a-judge paradigm for fully automated, objective, and general evaluation. Moreover, WebCoderBench adopts human-preference-aligned weights over metrics to yield interpretable overall scores. Experiments across 12 representative LLMs and 2 LLM-based agents show that there exists no dominant model across all evaluation metrics, offering an opportunity for LLM developers to optimize their models in a targeted manner for a more powerful version.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。