打造可无限交互的动态虚拟世界,支持多人实时沉浸体验。
Infinite Worlds with Versatile Interactions

- 采用因果预训练实现无边界交互,保持输出质量一致。
- 实现实时响应,可驱动720p/60fps视频流,延迟极低。
- 引入智能体协作机制,支持丰富动作与动态环境生成。
我们提出LingBot-World 2.0(又称LingBot-World-Infinity),对前代模型进行四方面升级:(1) 基于精心设计的因果预训练范式,实现无限制交互范围,同时维持输出质量稳定;(2) 从基础模型中蒸馏出实时变体,确保快速响应,足以支持720p、60帧每秒的视频流渲染;(3) 相较于旧版本,新增多样化的交互元素,包括攻击、射箭、施法、射击等更广泛动作,以及更多文本驱动事件;(4) 首次在世界建模领域集成智能体调度系统,由领航智能体规划角色行为,导演智能体随场景推进生成新环境元素。为支持共享体验,我们还开发了多玩家同步沉浸界面。主模型为140亿参数,搭配13亿参数轻量级版本,可在单张GPU上轻松部署。
原文摘要 · Abstract (English)
We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades. (1) Our model achieves an unbounded interaction horizon while maintaining consistent output quality, benefiting from a carefully crafted causal pretraining paradigm. (2) Through distilling a real-time variant from the base model, our system guarantees rapid response time, sufficient to drive 720p video streams at 60 fps. (3) Compared to the previous version, this update introduces highly diverse interactive elements, comprising a broader spectrum of actions (e.g., attacking, archery, spell-casting, and shooting) alongside a richer variety of text-driven events. (4) We pioneer the integration of an agentic harness within the domain of world modeling, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses. Additionally, to facilitate a shared experience, we develop an interface that permits multiple players to simultaneously immerse themselves in this vivid world simulator. We pair our primary 14B model with a lightweight 1.3B counterpart, which supports effortless deployment on a single GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。