实时交互式建模中实现长期几何一致性,兼顾速度与记忆保持。
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

- 用双动作表示响应用户键盘鼠标输入,实现稳定控制。
- 通过动态重建上下文和时间重映射,缓解长期记忆衰减。
- 提出上下文强制蒸馏法,支持模型在实时下使用长程信息。
本文提出WorldPlay,一种流式视频扩散模型,可在实时交互建模中实现长期几何一致性,突破当前方法在速度与内存之间的权衡瓶颈。其核心由三大要素构成:1)采用双动作表示,增强对用户键盘与鼠标输入的鲁棒性控制;2)提出重构上下文记忆机制,从过往帧动态重建上下文,并利用时间重映射使长期重要的几何帧持续可访问,有效缓解记忆衰减;3)设计一种面向记忆感知模型的新型蒸馏方法——上下文强制,通过教师-学生间上下文对齐,保持学生模型对长程信息的利用能力,在实现24帧/秒实时生成的同时防止误差累积。整体系统可在720p分辨率下生成长时间序列视频,性能优于现有方法,并在多样场景中展现强泛化能力。项目页面与在线演示详见:https://3d-models.hunyuan.tencent.com/world/ 及 https://3d.hunyuan.tencent.com/sceneTo3D。
原文摘要 · Abstract (English)
This paper presents WorldPlay, a streaming video diffusion model that enables real-time, interactive world modeling with long-term geometric consistency, resolving the trade-off between speed and memory that limits current methods. WorldPlay draws power from three key ingredients. 1) We use a Dual Action Representation to enable robust action control in response to the user's keyboard and mouse inputs. 2) To enforce long-term consistency, our Reconstituted Context Memory dynamically rebuilds context from past frames and uses temporal reframing to keep geometrically important but long-past frames accessible, effectively alleviating memory attenuation. 3) We also propose Context Forcing, a novel distillation method designed for memory-aware model. Aligning memory context between the teacher and student preserves the student's capacity to use long-range information, enabling real-time speeds while preventing error drift. Taken together, WorldPlay generates long-horizon streaming 720p video at 24 FPS with superior consistency, comparing favorably with existing techniques and showing strong generalization across diverse scenes. Project page and online demo can be found: https://3d-models.hunyuan.tencent.com/world/ and https://3d.hunyuan.tencent.com/sceneTo3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。