arXiv:2605.10523cs.CV2026-05中稿 · CVPR被引 1

通过语义对齐提升人物图像动画的结构稳定性和身份一致性

Improving Human Image Animation via Semantic Representation Alignment

论文配图:Improving Human Image Animation via Semantic Representation Alignment
图 1 · 摘自论文原文
  • 用结构与身份表示对齐替代传统像素监督,增强生成稳定性
  • 在长视频生成中显著减少肢体扭曲和面部失真,保持角色一致性
  • 适合需要高保真人像动画的应用,如虚拟偶像、影视特效

图像到视频生成领域已取得显著进展,但长期视频或剧烈动作下仍存在肢体扭曲和面部失真问题。现有方法依赖人体特定语义表示(如密集姿态或身份嵌入)作为条件输入,但会降低生成灵活性,且仅依赖RGB像素监督,忽视3D几何关系与时间连贯性。为此,本文提出SemanticREPA,将语义表示作为对齐监督信号。首先训练结构对齐模块,使视频潜在特征与深度估计特征对齐;固定后,用于指导扩散模型的结构表示,实现结构校正。同时设计身份对齐模块,对齐生成视频的身份表示与人脸识别特征,并利用预测结构优化相关区域的身份恢复。实验表明,该方法在复杂动作序列中显著提升生成质量与角色一致性。

原文摘要 · Abstract (English)

The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, especially when generating long videos or modeling intensive motions. Existing human image animation works address these issues by incorporating human-specific semantic representations, e.g., dense poses or ID embeddings, as additional conditions. However, conditioning on these representations could decrease the generation flexibility. Moreover, their reliance on RGB pixel supervision also lacks emphasis on learning necessary 3D geometric relationships and temporal coherence. In contrast, we introduce a novel approach named SemanticREPA that leverages these semantic representations as supervision signals through representation alignment. Specifically, we begin by training a structure alignment module that aligns the structure representations obtained from video latents with video depth estimation features. We then fix the pretrained module, and utilize it to provide additional supervision on the structure representations of the diffusion models, achieving structure rectification to generate coherent and stable human structures. Simultaneously, we develop an ID alignment module to align the ID representations of the generated videos to face recognition features. We further propose to use the predicted structure representations to refine identity restoration in relevant regions. With structure and ID alignment, our method demonstrates superior quality on extended character motions and enhanced character consistency.

图像动画扩散模型语义对齐身份一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。