用未来提示生成连贯教学视频,解决长序列动作一致性难题
SneakPeek: Future-Guided Instructional Streaming Video Generation
- 基于扩散模型的自回归框架,预判未来关键帧并引导生成
- 在复杂多步骤任务上生成语义准确、时序连贯的教学视频
- 适合需要交互式、分步控制的教育与内容创作场景
教学视频生成是一项新兴任务,旨在从文本描述中合成连贯的流程演示。该能力在内容创作、教育及人机交互中具有广泛应用前景,但现有视频扩散模型在长序列多步骤动作中难以保持时间一致性和可控性。本文提出SneakPeek,一种面向未来引导的流式教学视频生成流水线,是一种基于扩散模型的自回归框架,可基于初始图像和结构化文本提示生成精确、分步的教学视频。方法引入三项关键创新:(1)预测性因果适应,即因果模型学习下一帧预测并预判未来关键帧;(2)未来引导的自强迫机制结合双区域键值缓存,缓解推理时的暴露偏差;(3)多提示条件控制,实现对多步骤指令的细粒度流程控制。上述组件共同缓解时序漂移,保持运动一致性,并支持交互式生成——未来提示更新可动态影响正在进行的流式视频生成。实验表明,该方法生成的视频在时间上连贯且语义忠实,能准确遵循复杂的多步骤任务描述。
原文摘要 · Abstract (English)
Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad implications for content creation, education, and human-AI interaction, yet existing video diffusion models struggle to maintain temporal consistency and controllability across long sequences of multiple action steps. We introduce a pipeline for future-driven streaming instructional video generation, dubbed SneakPeek, a diffusion-based autoregressive framework designed to generate precise, stepwise instructional videos conditioned on an initial image and structured textual prompts. Our approach introduces three key innovations to enhance consistency and controllability: (1) predictive causal adaptation, where a causal model learns to perform next-frame prediction and anticipate future keyframes; (2) future-guided self-forcing with a dual-region KV caching scheme to address the exposure bias issue at inference time; (3) multi-prompt conditioning, which provides fine-grained and procedural control over multi-step instructions. Together, these components mitigate temporal drift, preserve motion consistency, and enable interactive video generation where future prompt updates dynamically influence ongoing streaming video generation. Experimental results demonstrate that our method produces temporally coherent and semantically faithful instructional videos that accurately follow complex, multi-step task descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。