arXiv:2507.16310cs.CV2025-07

让不同物体间动作迁移更自然,无需训练即可实现高保真运动转移。

MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

  • 通过语义特征匹配与形状重定向实现细粒度对应关系解析
  • 在外观和结构差异大的物体间仍保持动作连贯性
  • 无需训练,适合快速生成跨物体动作视频

现有文本到视频方法在参考物体与目标物体外观或结构差异较大时,难以平滑传递动作。为此,我们提出 MotionShot,一种无需训练的框架,能够精细解析参考物与目标物之间的对应关系,实现高保真动作迁移并保持外观一致性。具体而言,MotionShot 首先进行语义特征匹配以确保高层对齐,再通过参考物到目标物的形状重定向建立低层形态对齐。通过时序注意力编码动作,MotionShot 能够在显著外观与结构差异下实现连贯的动作跨物体迁移,实验验证了其有效性。项目页面见:https://motionshot.github.io/。

原文摘要 · Abstract (English)

Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a training-free framework capable of parsing reference-target correspondences in a fine-grained manner, thereby achieving high-fidelity motion transfer while preserving coherence in appearance. To be specific, MotionShot first performs semantic feature matching to ensure high-level alignments between the reference and target objects. It then further establishes low-level morphological alignments through reference-to-target shape retargeting. By encoding motion with temporal attention, our MotionShot can coherently transfer motion across objects, even in the presence of significant appearance and structure disparities, demonstrated by extensive experiments. The project page is available at: https://motionshot.github.io/.

文本生成视频动作迁移零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。