arXiv:2605.18601cs.CV2026-05被引 6

用自然语言控制多角色视频世界,实现跨角色灵活操作。

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

论文配图:Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
图 1 · 摘自论文原文
  • 以自然语言作为动作接口,实现每0.25秒的细粒度控制。
  • 跨角色迁移准确率达89%,远超基线43%;支持未见词汇指令。
  • 支持长时稳定生成,2小时滚动输出仍保持高质量,适合游戏与动画开发。

当前交互式视频世界模型虽视觉逼真,但在多实体精细控制和跨实体、跨世界泛化上存在不足。根源在于动作接口:传统协议(如动画编号、设备输入)在设计时绑定特定实体或引擎。本文提出以自然语言为接口,首次实现每帧(0.25秒)级的自然语言条件化,支持多实体同步控制及超越固定渲染管线的概念级跨实体迁移。结合预训练双向视频主干与帧级文本交叉注意力,通过ODE初始化的自强化蒸馏与RoPE解耦滑动KV缓存,实现实时长时序流式生成。在跨实体迁移任务中表现优异(89% vs. 43%),对未见词汇提示准确率达90%(基线0%)。2步学生模型在480p下维持19.7 FPS,2小时滚动输出稳定,FVD表现良好。同架构应用于《拳皇》游戏,仅调整实体动作词汇槽即可适配。已发布包含结构化动作元数据的《艾尔登法环》战斗片段预览集,完整数据将随项目公开。

原文摘要 · Abstract (English)

Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace this gap to the action interface: standard control protocols (e.g. animation IDs, device inputs, scene-level captions) bind action semantics to specific entities or engines at design time. We propose natural language as the interface to unlock expressiveness that no prior interface can achieve, and we present Incantation, the first interactive video world model with per-latent-frame (0.25 s) natural-language conditioning that supports simultaneous multi-entity control and concept-level cross-entity transfer beyond any fixed rendering pipeline. We pair a pretrained bidirectional video backbone with frame-local text cross-attention, and enable real-time long-horizon streaming through ODE-initialized Self-Forcing distillation with a RoPE-decoupled sliding KV-cache. We surpass the Action-Index baseline on cross-entity transfer (89% vs. 43%) and out-of-vocabulary prompts (90% vs. 0%), and our 2-step student sustains 19.7 FPS at 480p with stable FVD over 2-hour rollouts. We further apply the same architecture and training recipe to The King of Fighters, changing only the per-entity action vocabulary slots. We have released a preview subset of the Incantation dataset at https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes, containing manually collected Elden Ring player-boss combat clips with structured action-oriented metadata. Larger-scale Elden Ring and KOF data will be released with the full project.

视频生成自然语言控制多实体交互模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。