通过解耦身份与动作,实现精准可控的视频人脸生成。
Customizing Video Portraits via Identity-ActionDecoupling

- 将身份特征与动作信息分离建模,避免干扰
- 生成视频保持身份一致且表情动作贴合文本描述
- 无需微调,适合个性化视频创作场景
身份保真文本到视频生成(IPT2V)旨在从参考图像和文本描述中合成时序连贯的视频,同时保持主体身份并精细控制面部动态。尽管近期方法如ID-Animator和ConsisID仅在推理时注入身份特征,却忽略了面部嵌入中蕴含的非身份信息,导致面部动作单调或不准确,难以匹配提示。本文提出身份-动作解耦(IaD)框架,以及身份解耦损失和文本对齐损失两个新损失函数来解决该问题。无需任何特定主体微调,IaD生成的视频能(1)维持跨时间的身份一致性,(2)展现出丰富、可控制的表情与场景变化,且高度契合输入文本。
原文摘要 · Abstract (English)
Identity-Preserving Text-to-Video Generation (IPT2V) seeks to synthesize a temporally coherent video from a reference image and a textual description, while simultaneously preserving the subject's identity and allowing fine-grained control over facial dynamics. Although recent methods such as ID-Animator and ConsisID inject identity features only at inference time, they ignored the ID-irrelevant information contained in Facial embedding, leading to monotonous or inaccurate facial movements that poorly follow the prompt. We introduce Identity-Action Decoupling (IaD) framework as well as two loss function Identity Decoupling Loss and Text Alignment Loss to solve this problem. Without any subject-specific fine-tuning, IaD yields videos that (1) maintain cross-temporal identity consistency and (2) exhibit rich, controllable expressions and scene variations that closely match the input text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。