arXiv:2602.06028cs.CV2026-02被引 40

通过长上下文教师指导,实现20秒以上稳定视频生成。

Context Forcing: Consistent Autoregressive Video Generation with Long Context

  • 用能感知全局历史的长上下文教师监督学生模型
  • 生成视频上下文长度超20秒,是现有方法的2到10倍
  • 适合需要长期一致性视频生成的研究与应用

当前实时长视频生成多采用流式微调策略,由短上下文(无记忆)教师指导长上下文学生模型。这种结构导致师生不匹配:教师仅能访问5秒窗口,无法指导学生捕捉全局时间依赖,限制了学生上下文长度。为此,我们提出Context Forcing框架,通过长上下文教师训练长上下文学生,消除监督错配,实现鲁棒训练。为支持极端时长(如2分钟),引入上下文管理系统,构建慢-快记忆架构,显著降低视觉冗余。实验表明,该方法可实现超过20秒的有效上下文长度——比LongLive和Infinite-RoPE等先进方法长2至10倍。借助扩展上下文,Context Forcing在长时间生成中保持更优一致性,在多个长视频评估指标上超越现有基线。

原文摘要 · Abstract (English)

Recent approaches to real-time long video generation typically employ streaming tuning strategies, attempting to train a long-context student using a short-context (memoryless) teacher. In these frameworks, the student performs long rollouts but receives supervision from a teacher limited to short 5-second windows. This structural discrepancy creates a critical \textbf{student-teacher mismatch}: the teacher's inability to access long-term history prevents it from guiding the student on global temporal dependencies, effectively capping the student's context length. To resolve this, we propose \textbf{Context Forcing}, a novel framework that trains a long-context student via a long-context teacher. By ensuring the teacher is aware of the full generation history, we eliminate the supervision mismatch, enabling the robust training of models capable of long-term consistency. To make this computationally feasible for extreme durations (e.g., 2 minutes), we introduce a context management system that transforms the linearly growing context into a \textbf{Slow-Fast Memory} architecture, significantly reducing visual redundancy. Extensive results demonstrate that our method enables effective context lengths exceeding 20 seconds -- 2 to 10 times longer than state-of-the-art methods like LongLive and Infinite-RoPE. By leveraging this extended context, Context Forcing preserves superior consistency across long durations, surpassing state-of-the-art baselines on various long video evaluation metrics.

视频生成长上下文一致性扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。