Yan实现实时交互式视频生成,支持跨域风格融合与多粒度编辑。
Yan: Foundational Interactive Video Generation
- 用压缩3D-VAE与缓存去噪实现1080P/60FPS实时仿真
- 通过分层自回归提示注入游戏知识,生成可动作控制的无限视频流
- 分离模拟与渲染,支持文本驱动的多粒度内容编辑
我们提出Yan,一个面向交互式视频生成的基础性框架,覆盖从仿真、生成到编辑的全流程。Yan包含三个核心模块:AAA级仿真:设计高压缩低延迟的3D-VAE,结合基于KV缓存的滑动窗口去噪推理,实现实时1080P/60FPS交互仿真。多模态生成:引入分层自回归标题方法,将游戏特定知识注入开放域多模态视频扩散模型(VDMs),并转化为逐帧、动作可控、实时无限的交互式视频生成器。当文本与视觉提示来自不同领域时,模型展现出强泛化能力,能根据用户提示灵活融合与组合跨域风格与机制。多粒度编辑:提出混合模型,显式解耦交互机制仿真与视觉渲染,通过文本实现交互过程中的多粒度视频内容编辑。整体上,Yan整合各模块,推动交互式视频生成从孤立能力迈向全面的AI驱动交互创作范式,为下一代创意工具、媒体与娱乐铺平道路。项目页面:https://greatx3.github.io/Yan/
原文摘要 · Abstract (English)
We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three core modules. AAA-level Simulation: We design a highly-compressed, low-latency 3D-VAE coupled with a KV-cache-based shift-window denoising inference process, achieving real-time 1080P/60FPS interactive simulation. Multi-Modal Generation: We introduce a hierarchical autoregressive caption method that injects game-specific knowledge into open-domain multi-modal video diffusion models (VDMs), then transforming the VDM into a frame-wise, action-controllable, real-time infinite interactive video generator. Notably, when the textual and visual prompts are sourced from different domains, the model demonstrates strong generalization, allowing it to blend and compose the style and mechanics across domains flexibly according to user prompts. Multi-Granularity Editing: We propose a hybrid model that explicitly disentangles interactive mechanics simulation from visual rendering, enabling multi-granularity video content editing during interaction through text. Collectively, Yan offers an integration of these modules, pushing interactive video generation beyond isolated capabilities toward a comprehensive AI-driven interactive creation paradigm, paving the way for the next generation of creative tools, media, and entertainment. The project page is: https://greatx3.github.io/Yan/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。