arXiv:2412.01343cs.CV2024-12中稿 · ACM MM 2024, code …被引 14

用文本控制视频动作迁移,实现精准人体动作生成。

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

  • 通过多模态大模型扩展提示词,聚焦外观特征。
  • 引入外观注入模块,分离动作与外观建模。
  • 适合需要精细动作控制的视频生成场景。

现有预训练文本到视频(T2V)模型在生成基础运动或镜头移动方面表现优异,但在复杂人体动作生成上存在显著局限。当前方法通常在少量特定动作视频上微调模型,难以有效解耦参考视频中的动作与外观,削弱了动作模式建模能力。为此,我们提出MoTrans,一种定制化动作迁移方法,可在新场景中生成相似动作。具体而言,引入基于多模态大语言模型(MLLM)的重描述器,将初始提示扩展为更关注外观的描述;并设计外观注入模块,将视频帧的外观先验适配至动作建模过程。来自重描述提示和视频帧的互补多模态表示,促进外观建模并实现动作与外观解耦。此外,我们设计了针对特定动作的嵌入以进一步增强动作建模。实验表明,该方法能从单个或多个参考视频中有效学习特定动作模式,在定制化视频生成任务中优于现有方法。

原文摘要 · Abstract (English)

Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exhibit significant limitations when generating intricate, human-centric motions. Current efforts primarily focus on fine-tuning models on a small set of videos containing a specific motion. They often fail to effectively decouple motion and the appearance in the limited reference videos, thereby weakening the modeling capability of motion patterns. To this end, we propose MoTrans, a customized motion transfer method enabling video generation of similar motion in new context. Specifically, we introduce a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process. These complementary multimodal representations from recaptioned prompt and video frames promote the modeling of appearance and facilitate the decoupling of appearance and motion. In addition, we devise a motion-specific embedding for further enhancing the modeling of the specific motion. Experimental results demonstrate that our method effectively learns specific motion pattern from singular or multiple reference videos, performing favorably against existing methods in customized video generation.

动作迁移文本生成视频多模态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。