arXiv:2510.04097cs.AI2025-10被引 1

用强化学习提升网页代码生成质量,关键在布局风格一致性评估。

WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning

  • 基于渲染结果设计新评估指标,衡量布局与风格一致性。
  • 在45.1万真实网页上训练,生成效果超越现有方法。
  • 适合前端自动化、快速原型开发的研究者与工程师。

自动化将UI图像转换为网页代码是前端开发与快速原型设计的关键任务。多模态大语言模型(MLLM)的发展使网页界面到代码的生成成为可能,但现有基准数据集在多样性与评估可靠性方面仍显不足。为此,我们提出WebRenderBench,一个包含45.1万真实门户网站网页的大规模基准,具有更高的多样性、复杂性和真实性。我们进一步设计了一种新评估指标,从最终渲染页面出发,衡量布局与风格一致性。该方法不依赖视觉推理或易受噪声和不对称性影响的结构对比,实现了更高效、客观、可靠的UI质量评估。最后,我们引入自动化布局与风格检测代理(ALISA),将该指标作为强化学习的奖励信号,用于在非对称网页上优化生成训练。实验表明,ALISA显著提升了生成性能,在多个指标上达到当前最优水平。

原文摘要 · Abstract (English)

Automating the conversion of UI images into web code is a critical task for front-end development and rapid prototyping. Advances in multimodal large language models (MLLMs) have made WebUI-to-Code increasingly feasible, yet existing benchmarks remain limited in data diversity and evaluation reliability. To address these issues, we present WebRenderBench, a large-scale benchmark of 45.1k webpages collected from real-world portal sites, offering greater diversity, complexity, and realism than prior benchmarks. We further propose a novel evaluation metric that measures layout and style consistency from the final rendered pages. Unlike vision-based methods that rely on costly LLM reasoning or structure-based comparisons vulnerable to noise and asymmetry, our approach enables more efficient, objective, and reliable UI quality assessment. Finally, we introduce the Automated Layout and Style Inspection Agent (ALISA), which integrates this metric into reinforcement learning as a reward signal to enhance training on crawled asymmetric webpages. Experiments show that ALISA significantly boosts generation performance, achieving state-of-the-art results across multiple metrics.

网页生成强化学习评估指标多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。