arXiv:2411.17048cs.CV2024-11ICCV被引 31

用奖励机制实现高保真人像视频定制,不损失动作与语义

PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation

  • 用奖励机制替代图像重建,避免训练推理偏差
  • 引入身份一致与语义一致双奖励,保持动作和内容自然性
  • 仅需一张参考图就能稳定生成,适合个性化视频创作

当前文本到视频生成在通用视频合成上进展显著,但在特定身份的人像视频定制方面仍存在挑战。核心难题在于注入身份后如何保持高身份保真度,同时不破坏原始动作动态和语义连贯性。现有方法依赖在文生图模型上重构身份图像,与文生视频模型分布不一致,导致训练-推理差距,引发动态与语义退化。为此,我们提出PersonalVideo框架,采用合成视频上的混合奖励监督而非简单的图像重建目标。首先引入身份一致性奖励,有效注入参考身份且无训练-推理间隙;其次设计语义一致性奖励,对齐生成视频的语义分布与原始文生视频模型,保持其动作与语义跟随能力。通过非重构奖励训练,并结合模拟提示增强,提升多语义场景下的鲁棒性,即使仅用单张参考图像也能实现稳定生成。大量实验表明,该方法在保持高身份保真度的同时,充分保留原文生视频模型的生成质量,优于现有方法。

原文摘要 · Abstract (English)

The current text-to-video (T2V) generation has made significant progress in synthesizing realistic general videos, but it is still under-explored in identity-specific human video generation with customized ID images. The key challenge lies in maintaining high ID fidelity consistently while preserving the original motion dynamic and semantic following after the identity injection. Current video identity customization methods mainly rely on reconstructing given identity images on text-to-image models, which have a divergent distribution with the T2V model. This process introduces a tuning-inference gap, leading to dynamic and semantic degradation. To tackle this problem, we propose a novel framework, dubbed $\textbf{PersonalVideo}$, that applies a mixture of reward supervision on synthesized videos instead of the simple reconstruction objective on images. Specifically, we first incorporate identity consistency reward to effectively inject the reference's identity without the tuning-inference gap. Then we propose a novel semantic consistency reward to align the semantic distribution of the generated videos with the original T2V model, which preserves its dynamic and semantic following capability during the identity injection. With the non-reconstructive reward training, we further employ simulated prompt augmentation to reduce overfitting by supervising generated results in more semantic scenarios, gaining good robustness even with only a single reference image. Extensive experiments demonstrate our method's superiority in delivering high identity faithfulness while preserving the inherent video generation qualities of the original T2V model, outshining prior methods.

视频生成身份定制奖励机制文生视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。