让多物体视频生成更精准,通过运动对齐与语义约束实现可控动作迁移。
MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer
- 基于流的运动信号提供稳定运动先验,支持多对象动作迁移。
- 提出物体-描述对齐损失与接地评分,提升生成物体与描述的空间和语义一致性。
- 在定量、定性和人类评估中均优于现有方法,适合复杂场景视频生成任务。
动作迁移技术可通过将参考视频中的时序动态转移至新视频,实现基于目标描述的可控视频生成。然而,现有基于扩散变压器(DiT)的方法仅限于单对象视频,限制了真实世界多物体场景中的精细控制。本文提出MotionGrounder,首个支持多对象可控动作迁移的DiT框架。其提出的基于流的运动信号(FMS)为生成目标视频提供稳定运动先验;物体-描述对齐损失(OCAL)将物体描述锚定到对应空间区域。进一步设计物体接地评分(OGS),联合评估生成物体与源物体的空间对齐性及与目标描述的语义一致性。实验表明,MotionGrounder在定量、定性和人类评估中均持续优于近期基线方法。
原文摘要 · Abstract (English)
Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)-based methods are limited to single-object videos, restricting fine-grained control in real-world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT-based framework that firstly handles motion transfer with multi-object controllability. Our Flow-based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object-Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a new Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。