评测AI视频生成中提示词逆向还原能力,发现现有模型仍严重不足。
VI-Bench: Benchmarking Prompt Inversion from AIGC Videos

- 构建包含1610万条真实提示和900个真人验证视频的基准测试集
- 最强模型在提示还原评分上仅达0.632,多镜头复杂控制下性能骤降
- 揭示当前视觉语言模型难以将视觉理解转化为可复现的生成控制
视频生成技术的发展使基于提示的控制日益重要,提示决定了视频内容与表现形式,如视觉风格或摄像机行为。理解提示的可还原性对创意重用和编辑至关重要,也关乎提示泄露风险评估。然而现有视频理解基准无法衡量此能力:描述可见内容的字幕,未必能恢复生成所需的控制信息。为此,我们提出VI-Bench,一个基于1610万条真实用户提示和900个真人验证AIGC视频的基准。该基准涵盖三个逐步增加难度的场景:单次语义定位、风格与摄像机行为控制、多镜头组合逆向。评估五个生成关键维度:主体、动作、场景、风格、摄像机。我们在VI-Bench上评估了18个代表性视觉语言模型(含2个专有模型和16个开源模型),使用“逆向得分”衡量提示级对齐度与再生视频的视频级保真度。结果表明模型存在显著局限:即使最强模型得分仅0.632;随着控制需求增强和多镜头推理引入,性能急剧下降;模型常生成看似合理但再生视频与参考偏差较大的提示。这些发现表明,视频提示逆向是一项独立且未充分评估的能力,要求模型将视觉理解转化为可稳定重播的生成控制。
原文摘要 · Abstract (English)
Recent advances in video generation have made prompt-based control increasingly central to AIGC video generation. Prompts specify what a video should depict and how it should be represented, controlling factors such as visual style or camera behavior. Understanding this recoverability is important both for creative reuse and editing, and for assessing prompt leakage risks. However, existing video understanding benchmarks do not measure this capability: a caption may describe what is visible, but a replayable prompt must recover the generation-relevant controls needed to reproduce the video. To address this gap, we introduce VI-Bench, a benchmark built from 16.1 million real-user prompts and 900 human-verified AIGC videos. VI-Bench spans three progressively harder settings, namely single-shot semantic grounding, control over style and camera behavior, and multi-shot compositional inversion, and evaluates five generation-critical dimensions: subject, action, scene, style, and camera. We evaluate 18 representative VLMs, including 2 proprietary and 16 open-source models on VI-Bench, using an Inversion Score that measures prompt-level alignment with the original prompt and video-level fidelity of the regenerated video. The results reveal substantial limitations: even the strongest model achieves only 0.632 on Inversion Score, performance degrades sharply as samples require richer control and multi-shot reasoning, and models often produce plausible prompts whose regenerated videos deviate from the reference. These findings show that video prompt inversion is a distinct and under-evaluated capability requiring models to transform visual understanding into replay-stable generative control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。