arXiv:2608.09573cs.CV2026-08

用视频诊断网页生成质量,精准定位失败原因

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

论文配图:VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation
图 1 · 摘自论文原文
  • 基于操作视频生成细粒度诊断任务,定位四类失败
  • 覆盖1.7千个诊断问题,发现6338次真实生成错误
  • 适合评估网页生成模型行为可靠性,研究者必看

自然语言驱动的“氛围编码”可实现一键生成视觉丰富且交互性强的网页应用,但其质量评估仍滞后。现有方法多评分孤立产物或最终结果,难以揭示失败原因。我们提出VideoVIBE,一个视频基础诊断基准,将人工操作的网页录制转化为细粒度诊断任务。该基准包含约1.7K个诊断视频问答实例,源自6,338次经验证的生成网页失败,涵盖语义逻辑、视觉运动、结构时序和功能四类错误。诊断主要基于视频中的呈现与行为,网页源码作为补充上下文。我们进一步提出V2Lens,一种无需训练、基于证据的多智能体系统,通过针对性视觉与代码验证,挑战并选择性修正初始视频诊断。在十三个闭源与开源视频多模态大模型中,Gemini-2.5-Flash表现最强,得分为64.54;而V2Lens达到71.72,提升7.18分。结果表明,视频基础评估能超越孤立产物与聚合结果,提供行为忠实且诊断清晰的生成质量分析。

原文摘要 · Abstract (English)

Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace. Existing evaluations often score isolated artifacts or final task outcomes, offering limited evidence about which failures occur and why. We introduce VideoVIBE, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks. It contains approximately 1.7K diagnostic Video QA instances derived from 6,338 verified failures across generated webpages, spanning semantic-logical, visual-motion, structural-temporal, and functional failures. Diagnoses are grounded primarily in recorded presentation and behavior, with webpage source code used as complementary context. We further propose V2Lens, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-code verification. Across thirteen closed-source and open-weight Video MLLMs, Gemini-2.5-Flash is the strongest standalone model with a score of 64.54, while V2Lens reaches 71.72, an improvement of 7.18 points. Together, our results show that video-grounded evaluation can move beyond isolated artifacts and aggregate outcomes toward a behaviorally faithful and diagnostically informative account of generated application quality.

网页生成视频评估诊断基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。