用块扩散技术加速世界模型推理,实现高质量长视频生成。
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
- 采用块扩散架构,结合扩散与自回归优势,提升视频生成连贯性。
- 支持可变长度生成,实现分钟级高质量视频实时输出。
- 适合需要交互式世界模拟的研究者与开发者使用。
世界模型是代理型AI、具身AI和游戏等领域中的核心模拟器,能够生成长时、物理真实且可交互的高质量视频。通过扩展这些模型,有望在视觉感知、理解与推理方面涌现新能力,推动超越当前以大语言模型为中心的视觉基础模型的新范式。关键突破在于半自回归(块扩散)解码范式,该方法在每个块内采用扩散方式生成视频令牌,并依赖先前块进行条件化,从而实现更连贯稳定的视频序列。其核心优势在于重新引入类似大语言模型的键值缓存机制,克服了标准视频扩散模型的局限,支持高效、可变长度和高质量生成。因此,Inferix被专门设计为下一代推理引擎,通过优化半自回归解码流程,实现沉浸式世界合成。该系统区别于高并发场景(如vLLM或SGLang)和传统视频扩散模型(如xDiT)。Inferix进一步提供交互式视频流与性能分析功能,支持实时交互与动态建模。同时,无缝集成LV-Bench——一个针对分钟级视频生成场景的细粒度评估基准,实现高效基准测试。我们希望社区共同推进Inferix,促进世界模型研究发展。
原文摘要 · Abstract (English)
World models serve as core simulators for fields such as agentic AI, embodied AI, and gaming, capable of generating long, physically realistic, and interactive high-quality videos. Moreover, scaling these models could unlock emergent capabilities in visual perception, understanding, and reasoning, paving the way for a new paradigm that moves beyond current LLM-centric vision foundation models. A key breakthrough empowering them is the semi-autoregressive (block-diffusion) decoding paradigm, which merges the strengths of diffusion and autoregressive methods by generating video tokens in block-applying diffusion within each block while conditioning on previous ones, resulting in more coherent and stable video sequences. Crucially, it overcomes limitations of standard video diffusion by reintroducing LLM-style KV Cache management, enabling efficient, variable-length, and high-quality generation. Therefore, Inferix is specifically designed as a next-generation inference engine to enable immersive world synthesis through optimized semi-autoregressive decoding processes. This dedicated focus on world simulation distinctly sets it apart from systems engineered for high-concurrency scenarios (like vLLM or SGLang) and from classic video diffusion models (such as xDiTs). Inferix further enhances its offering with interactive video streaming and profiling, enabling real-time interaction and realistic simulation to accurately model world dynamics. Additionally, it supports efficient benchmarking through seamless integration of LV-Bench, a new fine-grained evaluation benchmark tailored for minute-long video generation scenarios. We hope the community will work together to advance Inferix and foster world model exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。