用一张图生成可键盘操控的动态交互世界,支持无限视频续播。
Yume: An Interactive World Generation Model
- 基于图像输入,通过相机运动量化与自回归视频生成实现动态世界构建。
- 支持无限时长视频生成,结合抗伪影机制与随机微分方程采样提升画质与控制精度。
- 适合游戏开发、虚拟现实及脑机接口方向研究者使用。
Yume旨在利用图像、文本或视频生成一个可交互、真实且动态的世界,支持通过外设或神经信号进行探索与控制。本文展示预览版 extit{method},其从输入图像生成动态世界,并允许用户通过键盘操作进行探索。为实现高保真、可交互的视频世界生成,我们设计了包含四个核心组件的框架:相机运动量化、视频生成架构、先进采样器与模型加速。首先,通过量化相机运动实现稳定训练与键盘输入的友好交互。其次,引入带记忆模块的掩码视频扩散变换器(MVDT),支持自回归方式下的无限视频生成。随后,在采样阶段引入无训练抗伪影机制(AAM)与基于随机微分方程的时间旅行采样(TTS-SDE),显著提升视觉质量与控制精度。此外,通过对抗性蒸馏与缓存机制协同优化实现模型加速。我们在高质量世界探索数据集 extit{sekai}上训练 extit{method},在多种场景与应用中表现优异。所有数据、代码与模型权重均已开源于https://github.com/stdstu12/YUME。Yume将每月更新以达成原始目标。项目页:https://stdstu12.github.io/YUME-Project/。
原文摘要 · Abstract (English)
Yume aims to use images, text, or videos to create an interactive, realistic, and dynamic world, which allows exploration and control using peripheral devices or neural signals. In this report, we present a preview version of \method, which creates a dynamic world from an input image and allows exploration of the world using keyboard actions. To achieve this high-fidelity and interactive video world generation, we introduce a well-designed framework, which consists of four main components, including camera motion quantization, video generation architecture, advanced sampler, and model acceleration. First, we quantize camera motions for stable training and user-friendly interaction using keyboard inputs. Then, we introduce the Masked Video Diffusion Transformer~(MVDT) with a memory module for infinite video generation in an autoregressive manner. After that, training-free Anti-Artifact Mechanism (AAM) and Time Travel Sampling based on Stochastic Differential Equations (TTS-SDE) are introduced to the sampler for better visual quality and more precise control. Moreover, we investigate model acceleration by synergistic optimization of adversarial distillation and caching mechanisms. We use the high-quality world exploration dataset \sekai to train \method, and it achieves remarkable results in diverse scenes and applications. All data, codebase, and model weights are available on https://github.com/stdstu12/YUME. Yume will update monthly to achieve its original goal. Project page: https://stdstu12.github.io/YUME-Project/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。