让游戏视频模型更快更准响应操作,支持低延迟交互与回放优化。
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

- 分阶段训练让模型在1-4步内精准生成动作控制视频。
- 在Minecraft中实现最高画质与最接近真实操作轨迹的运动一致性。
- 支持实时交互+回放重精炼双模式,兼顾速度与质量。
动作条件视频世界模型需实现低延迟因果生成并可靠响应游戏原生控制。尽管因果蒸馏可实现一至几步视频合成,但将其扩展至交互式世界模型仍具挑战,因离散键盘状态与连续鼠标运动必须与压缩的时序潜在块对齐。本文提出ForgeWM,通过领域适应、教师强制因果训练、因果一致性蒸馏及与双向教师的在线策略分布匹配,将双向动作条件视频生成器转化为高效少步世界模型。所生成的预算特化学生模型在稳态去噪预算为1、2、4步时表现优异。系统还支持双路径部署:低延迟交互与可选回放时间重精炼,其中一步学生可重新去噪并优化已保存草稿。在配对的Minecraft轨迹上,ForgeWM在图像质量、参考运动轮廓一致性、动作符号准确率和鼠标控制精度上均优于现有系统,且参考LPIPS最低;该四阶段方法同样适用于手柄控制的FPS游戏。回放重精炼达到四步参考质量,同时距离实际轨迹约三倍更近于从噪声再生。结果证明了ForgeWM在可控少步视频生成中的有效性。
原文摘要 · Abstract (English)
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。