构建可扩展网页生成评估框架,提升大模型生成网页的细节质量。
WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation
- 用智能爬虫持续抓取真实网页,扩充高质量训练数据。
- 按模块结构化组织网页内容、布局与视觉元素,支持精细监督。
- 分区块评估文本、布局和视觉一致性,适合网页生成研究者使用。
受大语言模型在编码与多模态理解方面进展的推动,我们提出 WebGen-V——一个面向指令到HTML生成的新基准与框架,旨在提升数据质量与评估粒度。WebGen-V包含三项关键创新:(1)无边界且可扩展的代理式爬虫框架,能持续采集真实网页,并用于增强现有基准;(2)结构化的分段数据表示,整合元数据、局部界面截图及JSON格式的文本与图像资源,明确对齐内容、布局与视觉组件,实现细粒度多模态监督;(3)分段级多模态评估协议,对齐文本、布局与视觉,实现高粒度评估。基于前沿大模型的实验与消融研究验证了结构化数据与分段评估的有效性,以及各组件的贡献。据我们所知,WebGen-V是首个实现指令到HTML生成中高粒度代理式爬取与评估的工作,提供从真实数据获取、网页生成到结构化多模态评估的统一流程。
原文摘要 · Abstract (English)
Witnessed by the recent advancements on leveraging LLM for coding and multimodal understanding, we present WebGen-V, a new benchmark and framework for instruction-to-HTML generation that enhances both data quality and evaluation granularity. WebGen-V contributes three key innovations: (1) an unbounded and extensible agentic crawling framework that continuously collects real-world webpages and can leveraged to augment existing benchmarks; (2) a structured, section-wise data representation that integrates metadata, localized UI screenshots, and JSON-formatted text and image assets, explicit alignment between content, layout, and visual components for detailed multimodal supervision; and (3) a section-level multimodal evaluation protocol aligning text, layout, and visuals for high-granularity assessment. Experiments with state-of-the-art LLMs and ablation studies validate the effectiveness of our structured data and section-wise evaluation, as well as the contribution of each component. To the best of our knowledge, WebGen-V is the first work to enable high-granularity agentic crawling and evaluation for instruction-to-HTML generation, providing a unified pipeline from real-world data acquisition and webpage generation to structured multimodal assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。