arXiv:2512.22096cs.CV2025-12被引 64

用文本或图片生成可实时交互的连续世界,速度快且支持内容控制。

Yume-1.5: A Text-Controlled Interactive World Generation Model

  • 结合上下文压缩与线性注意力,实现长视频持续生成。
  • 通过双向注意力蒸馏加速推理,支持实时流式渲染。
  • 文本控制事件生成,适合游戏开发与虚拟场景构建。

近期方法已展示利用扩散模型生成可交互、可探索世界的潜力。然而,多数方法存在参数量过大、推理步骤冗长、历史上下文急剧增长等问题,严重限制实时性能,且缺乏文本控制生成能力。为此,我们提出 extit{Yume-1.5},一种从单张图像或文本提示生成真实、可交互、连续世界的新框架。该框架支持键盘操作探索生成世界,包含三个核心组件:(1) 集成统一上下文压缩与线性注意力的长视频生成框架;(2) 基于双向注意力蒸馏与增强文本嵌入策略的实时流式加速方案;(3) 文本控制的世界事件生成方法。代码已提供于补充材料中。

原文摘要 · Abstract (English)

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy inference steps, and rapidly growing historical context, which severely limit real-time performance and lack text-controlled generation capabilities. To address these challenges, we propose \method, a novel framework designed to generate realistic, interactive, and continuous worlds from a single image or text prompt. \method achieves this through a carefully designed framework that supports keyboard-based exploration of the generated worlds. The framework comprises three core components: (1) a long-video generation framework integrating unified context compression with linear attention; (2) a real-time streaming acceleration strategy powered by bidirectional attention distillation and an enhanced text embedding scheme; (3) a text-controlled method for generating world events. We have provided the codebase in the supplementary material.

世界生成扩散模型实时交互文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。