通过阴影对学习统一动态表征,实现视频世界模型的任意动作精准控制。
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

- 利用阴影对分离动作与外观,构建跨场景可复用的动作表征
- 跨阴影预测使模型自动剥离外观差异,保留核心动作信息
- 无需标签或微调,可在新场景中直接重放演示动作
我们提出ShadowDancer,一种实现交互式视频世界模型任意动作、帧级控制的新方法。现有接口要么松散编码动作,让模型自行推演;要么依赖结构化信号,仅适用于特定动作族且难以获取,导致跨动态的精确控制仍不现实。示范视频虽能逐帧指定动态,但仅呈现特定外观下的单一阴影,导致动作迁移效果差。ShadowDancer通过两大创新解决:(1) 阴影对——由我们的阴影库大规模构建,即同一动态在独立重采样外观下生成的视频对,使动作家族可被精确控制的前提得以满足;(2) 跨阴影预测——通过从一个阴影预测另一个,自动舍弃重采样内容,保留共同动态信息,从而获得统一动态表征,驱动块因果世界模型。任何演示片段均可作为可重用动作资产,在无标签、无运动估计、无微调情况下于新环境重播。实验显示,在多种动态族上优于强基准的潜在动作与交互世界模型,平均盲测胜率86%。
原文摘要 · Abstract (English)
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。