用视觉语言模型定位视频生成中的细微错误,提升评估精度。
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
- 提出新任务Spotlight,精准定位视频生成中的局部错误。
- 发现物理和提示遵循性错误更普遍且持续时间长。
- 验证当前VLM在错误识别上远不如人类,提出优化策略提升近2倍性能。
当前文本到视频模型(T2V)虽能生成高质量、时序连贯且视觉逼真的视频,但仍存在细微且局部的错误。现有评估方法多为整体评价,难以定位具体错误发生的时间与类型。为此,我们提出Spotlight任务,旨在定位并解释视频生成错误。基于200个多样化文本提示和三种先进视频生成器(Veo 3、Seedance、LTX-2),生成600个视频,并标注超过1600个细粒度错误,涵盖运动、物理、提示遵循等六类。结果显示,提示遵循性和物理错误更为普遍,且持续时间较长;而外观出现/消失及身体姿态错误则多出现在短片段中。我们评估了当前视觉语言模型(VLMs)在Spotlight任务上的表现,发现其错误识别与定位能力显著落后于人类。通过引入推理时策略,性能提升近2倍。该任务为构建细粒度评估工具和更复杂的奖励模型开辟了新路径。
原文摘要 · Abstract (English)
Current text-to-video models (T2V) can generate high-quality, temporally coherent, and visually realistic videos. Nonetheless, errors still often occur, and are more nuanced and local compared to the previous generation of T2V models. While current evaluation paradigms assess video models across diverse dimensions, they typically evaluate videos holistically without identifying when specific errors occur or describing their nature. We address this gap by introducing Spotlight, a novel task aimed at localizing and explaining video-generation errors. We generate 600 videos using 200 diverse textual prompts and three state-of-the-art video generators (Veo 3, Seedance, and LTX-2), and annotate over 1600 fine-grained errors across six types, including motion, physics, and prompt adherence. We observe that adherence and physics errors are predominant and persist across longer segments, whereas appearance-disappearance and body pose errors manifest in shorter segments. We then evaluate current VLMs on Spotlight and find that VLMs lag significantly behind humans in error identification and localization in videos. We propose inference-time strategies to probe the limits of current VLMs on our task, improving performance by nearly 2x. Our task paves a way forward to building fine-grained evaluation tools and more sophisticated reward models for video generators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。