用文本/图像一键生成可交互的持续演化的虚拟世界
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

- 基于150亿参数扩散模型,按镜头轨迹自回归生成视频块
- 支持24帧/秒、540p/720p画面,长时生成误差低
- 适合研究交互式世界建模与开放平台构建的开发者
与传统游戏开发依赖繁琐资源制作不同,视频世界模型能根据用户输入即时生成可交互环境。该技术可从文本、图像或视频创建定制化、可探索且持续演进的虚拟世界。实现这一目标需具备四大能力:交互性、时空持久一致性、稳定长时生成和高效响应。我们提出AlayaWorld,一个交互式长时视频世界模型,可生成24帧/秒、540p及720p分辨率的视频。其基于150亿参数的视频扩散变换器,沿相机轨迹自回归生成短时潜空间块,并支持可切换的文本提示。模型采用有界视觉上下文,结合持久性参考帧、压缩时间历史、对齐几何的空间记忆及近期帧条件。为减少长期漂移,训练时使用自身回放产生的损坏历史与预测残差。我们进一步引入离散自回归蒸馏框架,融合分布匹配蒸馏、自强制++与一致性蒸馏,将推理步骤从约30步降至每块4步。在iWorld-Bench评测中,AlayaWorld在长时生成任务上表现最优。AlayaWorld被设计为全栈开源项目,旨在为未来交互式视频世界模型研究提供可扩展基础。
原文摘要 · Abstract (English)
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and efficient response. We present AlayaWorld, an interactive long-horizon video world model that generates 24-fps video at 540p and 720p. Built on a 15B video diffusion transformer, AlayaWorld generates short latent chunks autoregressively under camera trajectories and switchable text prompts. Its bounded visual context combines a persistent sink frame, compressed temporal history, geometry-aligned spatial memory, and recent-frame conditioning. To reduce long-term drift, the model is trained with corrupted histories and prediction residuals collected from its own roll-outs. We further introduce a discrete autoregressive distillation formulation that combines distribution-matching distillation, self-forcing++, and consistency distillation, reducing inference from approximately 30 sampling steps to four steps per chunk. On iWorld-Bench, AlayaWorld achieves the best performance over long-horizon generation. Conceived as a full-stack, open-source, and long-term project, AlayaWorld is intended to provide an extensible foundation for future research on interactive video world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。