用离散动作指令提升视频世界模型的可控性与稳定性
DisCo: World Models with Discrete Camera Motion Control

- 采用离散动作基元替代连续轨迹,增强动作可区分性
- 在复杂运动序列中实现更可靠的指令跟随,视觉质量保持良好
- 适合需要精准交互控制的视频生成场景
可控视频世界模型旨在实现交互式世界探索,要求模型在执行明确动作指令时保持视觉质量与时间连贯性。然而,现有方法多依赖连续相机轨迹作为动作条件,导致在复杂运动序列下动作跟随不可靠。本文指出动作表示纠缠是关键瓶颈,连续相机表征在不同运动模式间特征相似度高,降低可控性。为此提出DisCo,一种基于紧凑离散动作基元的可控视频世界模型,提升动作可分性。同时构建DisCoBench,一个涵盖短期、长时程及高度动态探索场景的综合评估基准。大量实验表明,DisCo在保持视觉质量的同时,显著提升动作跟随可靠性。
原文摘要 · Abstract (English)
Controllable video world models target interactive world exploration, where models must faithfully execute explicit action commands while preserving visual quality and temporal coherence. However, most existing approaches rely on continuous camera trajectories as action conditions, which often lead to unreliable action following, especially under complex motion sequences. In this work, we identify action representation entanglement as a key bottleneck in controllable video generation, and show that continuous camera representations lead to high feature similarity across distinct motion patterns, degrading action controllability. Based on this insight, we propose DisCo, a controllable video world model that conditions generation on a compact set of discrete action primitives to improve action separability. We further introduce DisCoBench, a comprehensive benchmark for evaluating the ability of models in short-term, long-horizon, and highly dynamic exploration scenarios. Extensive experiments demonstrate that DisCo achieves significantly more reliable action following while preserving visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。