arXiv:2503.10096cs.CV2025-03被引 1

提出一种自监督的运动表征,实现高效且逼真的肖像视频生成。

A Self-supervised Motion Representation for Portrait Video Generation

  • 用自监督掩码运动编码器将动作压缩为紧凑的1D隐状态。
  • 基于音频信号生成运动序列,实现81%真实感胜率超越当前最佳模型。
  • 适合需要高效推理和高保真视频生成的研究与应用。

近期肖像视频生成技术进展显著,但现有方法高度依赖人类先验或预训练生成模型,基于人类先验的运动表征易引入不自然动作,而依赖预训练模型的方法则常存在推理效率低的问题。为此,本文提出语义隐式运动(Semantic Latent Motion, SeMo),一种紧凑且富有表现力的运动表征。基于该表征,方法在保证高质量视觉效果的同时实现高效推理。SeMo采用三步框架:抽象、推理与生成。首先,在抽象阶段,使用精心设计的掩码运动编码器,通过自监督学习将主体运动状态压缩为一个紧凑的1维隐运动(1D token);其次,在推理阶段,基于驱动音频信号高效生成运动序列;最后,在生成阶段,运动动态作为条件信息引导运动解码器,从参考帧合成逼真的目标视频过渡。得益于SeMo的紧凑性和表达能力,本方法实现了高效的运动表征与高质量视频生成。用户研究显示,其在真实感上以81%胜率超越当前最先进模型。大量实验进一步验证了其强大的压缩能力、重建质量与生成潜力。

原文摘要 · Abstract (English)

Recent advancements in portrait video generation have been noteworthy. However, existing methods rely heavily on human priors and pre-trained generative models, Motion representations based on human priors may introduce unrealistic motion, while methods relying on pre-trained generative models often suffer from inefficient inference. To address these challenges, we propose Semantic Latent Motion (SeMo), a compact and expressive motion representation. Leveraging this representation, our approach achieve both high-quality visual results and efficient inference. SeMo follows an effective three-step framework: Abstraction, Reasoning, and Generation. First, in the Abstraction step, we use a carefully designed Masked Motion Encoder, which leverages a self-supervised learning paradigm to compress the subject's motion state into a compact and abstract latent motion (1D token). Second, in the Reasoning step, we efficiently generate motion sequences based on the driving audio signal. Finally, in the Generation step, the motion dynamics serve as conditional information to guide the motion decoder in synthesizing realistic transitions from reference frame to target video. Thanks to the compact and expressive nature of Semantic Latent Motion, our method achieves efficient motion representation and high-quality video generation. User studies demonstrate that our approach surpasses state-of-the-art models with an 81% win rate in realism. Extensive experiments further highlight its strong compression capability, reconstruction quality, and generative potential.

视频生成自监督运动表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。