arXiv:2604.10030cs.CV2026-04被引 5

让视频生成按时间顺序精准呈现多个事件,避免概念混淆。

Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation

  • 推理时通过注意力惩罚机制,分段控制每段视频的提示词
  • 显著提升多事件视频的时间对齐度与视觉质量
  • 无需修改模型结构,适合电影级视频生成场景

视频扩散模型在生成高质量视频方面取得显著进展,但难以准确表达真实视频中多个事件的时间先后顺序,且缺乏对语义概念出现时间、持续时长及事件顺序的显式控制。这对电影级视频合成尤为重要,因为连贯叙事依赖精确的时间安排和事件转换。使用单一段落提示描述复杂事件序列时,模型常出现语义纠缠,导致不同时间点的概念相互干扰,影响文本-视频对齐。为此,我们提出 Prompt Relay,一种推理时的即插即用方法,实现多事件视频生成中的细粒度时间控制,无需架构修改和额外计算开销。该方法在交叉注意力机制中引入惩罚项,使每个时间片段仅关注其对应的提示,从而实现逐时序表达单一语义概念,改善时间提示对齐,减少语义干扰,提升视觉质量。

原文摘要 · Abstract (English)

Video diffusion models have achieved remarkable progress in generating high-quality videos. However, these models struggle to represent the temporal succession of multiple events in real-world videos and lack explicit mechanisms to control when semantic concepts appear, how long they persist, and the order in which multiple events occur. Such control is especially important for movie-grade video synthesis, where coherent storytelling depends on precise timing, duration, and transitions between events. When using a single paragraph-style prompt to describe a sequence of complex events, models often exhibit semantic entanglement, where concepts intended for different moments in the video bleed into one another, resulting in poor text-video alignment. To address these limitations, we propose Prompt Relay, an inference-time, plug-and-play method to enable fine-grained temporal control in multi-event video generation, requiring no architectural modifications and no additional computational overhead. Prompt Relay introduces a penalty into the cross-attention mechanism, so that each temporal segment attends only to its assigned prompt, allowing the model to represent one semantic concept at a time and thereby improving temporal prompt alignment, reducing semantic interference, and enhancing visual quality.

视频生成扩散模型时间控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。