arXiv:2411.16718cs.CVcs.AI2024-11CVPR被引 13

用形式化验证评估文本生成视频的时序一致性,更贴近真实应用需求。

Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification

  • 将提示转为时序逻辑规范,视频转为自动机,严格验证对齐性。
  • 相比现有指标,与人工评估相关性提升5倍以上。
  • 适合关注安全关键场景中视频生成准确性的研究者。

Sora、Gen-3、MovieGen和CogVideoX等文本到视频模型的进展正推动合成视频生成边界,已在机器人、自动驾驶和娱乐领域得到应用。尽管已有多种评估指标和基准,但这些指标侧重视觉质量和流畅性,忽视了对安全关键应用至关重要的时序保真度和文本-视频对齐性。为此,我们提出NeuS-V,一种基于神经符号形式化验证技术的新型合成视频评估指标。该方法将提示转换为形式化的时序逻辑(TL)规范,并将生成视频转化为自动机表示,通过形式化检查视频自动机是否满足TL规范来评估对齐性。此外,我们构建了一个包含时序扩展提示的数据集,用于评估主流视频生成模型。结果表明,NeuS-V与人工评估的相关性比现有指标高出5倍以上。我们的评估还发现,当前模型在处理这类时序复杂提示时表现不佳,凸显了未来改进文本到视频生成能力的必要性。

原文摘要 · Abstract (English)

Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these models become prevalent, various metrics and benchmarks have emerged to evaluate the quality of the generated videos. However, these metrics emphasize visual quality and smoothness, neglecting temporal fidelity and text-to-video alignment, which are crucial for safety-critical applications. To address this gap, we introduce NeuS-V, a novel synthetic video evaluation metric that rigorously assesses text-to-video alignment using neuro-symbolic formal verification techniques. Our approach first converts the prompt into a formally defined Temporal Logic (TL) specification and translates the generated video into an automaton representation. Then, it evaluates the text-to-video alignment by formally checking the video automaton against the TL specification. Furthermore, we present a dataset of temporally extended prompts to evaluate state-of-the-art video generation models against our benchmark. We find that NeuS-V demonstrates a higher correlation by over 5x with human evaluations when compared to existing metrics. Our evaluation further reveals that current video generation models perform poorly on these temporally complex prompts, highlighting the need for future work in improving text-to-video generation capabilities.

文本生成视频形式化验证时序对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。