通过融合场景与动作提示,实现长视频连贯生成。
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
- 用动态启发的提示混合策略,双向加权融合时序片段。
- 在不额外训练下,生成时序一致且视觉吸引的长视频。
- 适合需要连贯视频叙事的创作者或内容生成应用。
从离散文本提示生成连贯的长视频序列仍具挑战,主要源于难以维持时间连贯性、语义一致性及场景-动作连续性。本文提出一种新型叙事框架,通过受动态启发的提示混合机制整合场景与动作提示。方法包含三部分:(i) 双向时间加权潜在空间混合策略,强化相邻视频段间的时间一致性;(ii) 动态启发的提示加权(DIPW)机制,基于CLIP对齐、叙事进展和时间平滑性,在每个扩散步自适应平衡场景与动作提示权重;(iii) 语义动作表示,编码高层动作语义以根据动作相似性调节过渡。潜在空间混合保持场景内空间一致性,时间加权混合引入双向时间约束,防止突兀切换。实验表明,该方法显著优于基线,在无需额外训练条件下生成时序一致且视觉生动的长视频,有效弥合短片段与扩展文本驱动视频叙事之间的差距。
原文摘要 · Abstract (English)
Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence, semantic consistency, and scene-action continuity across segments. We propose a novel storytelling framework that integrates scene and action prompts through dynamics-inspired prompt mixing. Our approach combines three key components: (i) a bidirectional time-weighted latent blending strategy that enforces temporal consistency between consecutive video segments, (ii) a dynamics-informed prompt weighting (DIPW) mechanism that adaptively balances scene and action prompts at each diffusion timestep based on CLIP-based alignment, narrative progression, and temporal smoothness, and (iii) a semantic action representation that encodes high-level action semantics to modulate transitions according to action similarity. Latent-space blending preserves spatial coherence within scenes, while time-weighted blending introduces bidirectional temporal constraints to prevent abrupt transitions. Together, these components enable fluid and coherent video narratives that faithfully reflect both scene context and action dynamics. Extensive experiments demonstrate that our method significantly outperforms baselines, producing temporally consistent and visually compelling long-form videos without any additional training, thereby bridging the gap between short clips and extended text-driven video storytelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。