arXiv:2603.13391cs.CV2026-03被引 5

用视频生成网页,构建首个专用评测基准

WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics

  • 基于视频演示重建网页,引入细粒度视觉评估标准
  • 175个网页数据集,自动评估与人工偏好一致率达96%
  • 适合研究视频到网页生成、多模态模型评估的学者

现有网页生成评测依赖文本提示或静态截图,但视频能更完整传递交互流程、切换时序和运动连续性,对真实网页重建至关重要。然而,视频驱动的网页生成仍缺乏专门评测基准。为此,我们提出WebVR,一个评估多模态大模型从演示视频中忠实重建网页能力的基准。WebVR包含175个跨类别网页,均通过受控合成流程构建,避免与现有网页重叠,确保演示多样性和真实性。我们设计了细粒度的人类对齐视觉评估标准,从多个维度评价生成结果。在19个模型上的实验表明,模型在精细风格和运动质量重建上仍有显著差距;基于该标准的自动评估与人类偏好一致性达96%。我们开源数据集、评估工具和基线结果,以推动视频到网页生成研究。

原文摘要 · Abstract (English)

Existing web-generation benchmarks rely on text prompts or static screenshots as input. However, videos naturally convey richer signals such as interaction flow, transition timing, and motion continuity, which are essential for faithful webpage recreation. Despite this potential, video-conditioned webpage generation remains largely unexplored, with no dedicated benchmark for this task. To fill this gap, we introduce WebVR, a benchmark that evaluates whether MLLMs can faithfully recreate webpages from demonstration videos. WebVR contains 175 webpages across diverse categories, all constructed through a controlled synthesis pipeline rather than web crawling, ensuring varied and realistic demonstrations without overlap with existing online pages. We also design a fine-grained, human-aligned visual rubric that evaluates the generated webpages across multiple dimensions. Experiments on 19 models reveal substantial gaps in recreating fine-grained style and motion quality, while the rubric-based automatic evaluation achieves 96% agreement with human preferences. We release the dataset, evaluation toolkit, and baseline results to support future research on video-to-webpage generation.

视频生成多模态评测基准网页重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。