arXiv:2608.21424cs.CVcs.GR2026-08

一个统一框架,让交互式视频生成与编辑更高效流畅。

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

论文配图:EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing
图 1 · 摘自论文原文
  • 基于DiT架构,用条件控制统一多种视频任务
  • 两阶段蒸馏提升少步生成质量与时间稳定性
  • 适合创意设计者快速迭代视频内容

交互式视频生成与编辑在创意设计中日益重要。本文提出EditStream:一种统一的视频生成与编辑框架。该框架基于DiT模型,通过灵活的任务特定条件,整合文本到视频、图像到视频、视频到视频、编辑传播、参考引导编辑和相机位姿变化等多种任务,实现单一系统内的灵活控制。为支持交互使用,采用两阶段蒸馏方法,结合速度矩匹配(VMM)与自回归展开:VMM在学生模型中间状态匹配条件速度矩,保持生成质量与运动一致性;展开则让学生暴露于自身自回归预测,增强时序稳定性。该方法有效缓解了少步自回归视频生成中的过饱和、运动退化、时序不稳定和复杂训练等问题。EditStream为高质量扩散模型与交互式创作流程提供了实用且可扩展的桥梁。

原文摘要 · Abstract (English)

Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible task-specific conditioning, and further transforms it into a fast, few-step autoregressive model for efficient streaming. It supports Text-to-Video, Image-to-Video, Video-to-Video, Editing Propagation, Reference-guided Video Editing, and Camera Pose Change, enabling flexible control over video generation, transformation, and editing within one system. To make the unified model practical for interactive use, we develop a two-stage distillation approach that combines Velocity Moment Matching (VMM) with autoregressive unrolling. VMM matches conditional velocity moments at student-reached intermediate states to preserve generation quality and motion, while unrolling exposes the student to its own autoregressive predictions to improve temporal stability. Together, they alleviate common challenges in few-step autoregressive video generation, including over-saturation, degraded motion, temporal instability, and complex training. EditStream provides a practical and scalable solution that bridges high-quality diffusion-based video models with interactive creative workflows.

视频生成交互编辑扩散模型自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。