分离运动与内容,实现跨类别动作迁移。
DisMo: Disentangled Motion Representations for Open-World Motion Transfer
- 从原始视频中学习独立于外观的抽象运动表示。
- 可在不同类别间无对应关系下完成动作迁移,保持动作准确性和提示一致性。
- 适合作品创作者、动画师及视频生成研究者使用。
文本到视频(T2V)和图像到视频(I2V)模型的进步,使得仅凭文字描述或初始帧即可生成视觉生动的动态视频。然而,这些模型通常无法提供显式的运动分离表示,限制了内容创作者的应用。为此,我们提出DisMo,一种通过图像空间重建目标直接从原始视频数据中学习抽象运动表示的新范式。该表示具有通用性,独立于外观、物体身份或姿态等静态信息,支持开放世界动作迁移,可在语义无关的实体间实现动作转移,无需对象对应关系,甚至在截然不同的类别之间也可行。与以往方法相比,我们的方法避免了运动保真度与提示遵循性的权衡,防止过拟合源结构或偏离描述动作。此外,该运动表示可与任意现有视频生成器结合使用轻量级适配器,轻松利用未来视频模型的进展。我们在多种动作迁移任务中验证了方法的有效性,并进一步证明所学表示适用于下游动作理解任务,在Something-Something v2和Jester基准上的零样本动作分类中持续优于当前最优的视频表示模型如V-JEPA。
原文摘要 · Abstract (English)
Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an explicit representation of motion separate from content, limiting their applicability for content creators. To address this gap, we propose DisMo, a novel paradigm for learning abstract motion representations directly from raw video data via an image-space reconstruction objective. Our representation is generic and independent of static information such as appearance, object identity, or pose. This enables open-world motion transfer, allowing motion to be transferred across semantically unrelated entities without requiring object correspondences, even between vastly different categories. Unlike prior methods, which trade off motion fidelity and prompt adherence, are overfitting to source structure or drifting from the described action, our approach disentangles motion semantics from appearance, enabling accurate transfer and faithful conditioning. Furthermore, our motion representation can be combined with any existing video generator via lightweight adapters, allowing us to effortlessly benefit from future advancements in video models. We demonstrate the effectiveness of our method through a diverse set of motion transfer tasks. Finally, we show that the learned representations are well-suited for downstream motion understanding tasks, consistently outperforming state-of-the-art video representation models such as V-JEPA in zero-shot action classification on benchmarks including Something-Something v2 and Jester. Project page: https://compvis.github.io/DisMo
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。