用神经符号反馈自动修复文本生成视频的逻辑错误,无需重新训练。
We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback
- 通过形式化视频表示分析生成错误反馈,定位不一致事件与对象。
- 在多种模型上使视频时序与逻辑对齐度提升近40%。
- 适合关注视频生成质量但无法微调模型的研究者和开发者。
当前文本到视频(T2V)生成模型因能从文本提示生成连贯视频而日益流行。然而,面对涉及多个物体或顺序事件的长而复杂的提示时,这些模型常难以生成语义和时序一致的视频。此外,训练或微调带来的高计算成本使直接改进变得不切实际。为此,我们提出NeuS-E,一种新颖的零训练视频优化流程,利用神经符号反馈自动提升视频生成质量,实现与提示的更好对齐。该方法首先通过分析形式化视频表示,生成神经符号反馈,精确定位语义不一致的事件、对象及其对应帧。随后,该反馈引导对原始视频进行针对性编辑。在开源及专有T2V模型上的大量实证评估表明,NeuS-E显著提升了多样化提示下的时序与逻辑一致性,提升幅度接近40%。
原文摘要 · Abstract (English)
Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to generate semantically and temporally consistent videos when dealing with longer, more complex prompts involving multiple objects or sequential events. Additionally, the high computational costs associated with training or fine-tuning make direct improvements impractical. To overcome these limitations, we introduce NeuS-E, a novel zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation, achieving superior alignment with the prompts. Our approach first derives the neuro-symbolic feedback by analyzing a formal video representation and pinpoints semantically inconsistent events, objects, and their corresponding frames. This feedback then guides targeted edits to the original video. Extensive empirical evaluations on both open-source and proprietary T2V models demonstrate that NeuS-E significantly enhances temporal and logical alignment across diverse prompts by almost 40%
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。