通过融合文本语义特征,生成更自然的人像视频。
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
- 用文本与图像融合提取高级语义身份特征,避开低层干扰
- 两阶段训练平衡浅层特征与高层特征,提升生成质量
- 适合需要高保真人像视频生成的场景,如影视制作
近期视频生成技术在身份保持视频生成(IPT2V)中取得显著进展,但现有方法仍存在‘复制粘贴’伪影和相似度低的问题,主要源于对低层人脸图像信息的过度依赖,导致面部僵硬且引入无关细节。为此,我们提出EchoVideo,采用两项关键策略:(1) 身份图像-文本融合模块(IITF),融合文本中的高层语义特征,捕捉干净的人脸身份表征,同时去除遮挡、姿态和光照变化的影响,避免伪影引入;(2) 两阶段训练策略,在第二阶段引入随机性,有选择地使用浅层人脸信息,以平衡浅层特征带来的保真度提升,同时减轻对其过度依赖。该策略促使模型在训练中更多利用高层特征,最终形成更鲁棒的人脸身份表示。实验表明,EchoVideo能有效保持人脸身份一致性并维持全身完整性,在生成高质量、可控性强、保真度高的视频方面表现优异。
原文摘要 · Abstract (English)
Recent advancements in video generation have significantly impacted various downstream applications, particularly in identity-preserving video generation (IPT2V). However, existing methods struggle with "copy-paste" artifacts and low similarity issues, primarily due to their reliance on low-level facial image information. This dependence can result in rigid facial appearances and artifacts reflecting irrelevant details. To address these challenges, we propose EchoVideo, which employs two key strategies: (1) an Identity Image-Text Fusion Module (IITF) that integrates high-level semantic features from text, capturing clean facial identity representations while discarding occlusions, poses, and lighting variations to avoid the introduction of artifacts; (2) a two-stage training strategy, incorporating a stochastic method in the second phase to randomly utilize shallow facial information. The objective is to balance the enhancements in fidelity provided by shallow features while mitigating excessive reliance on them. This strategy encourages the model to utilize high-level features during training, ultimately fostering a more robust representation of facial identities. EchoVideo effectively preserves facial identities and maintains full-body integrity. Extensive experiments demonstrate that it achieves excellent results in generating high-quality, controllability and fidelity videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。