生成带精细表情的个性化虚拟形象,确保表情多样且身份一致
Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
- 用多模态扩散Transformer,结合身份与表情特征生成图像
- 在多种细微表情下保持身份一致性,优于现有方法
- 适合虚拟社交、游戏、数字人等需要真实表情的应用
2D虚拟形象广泛应用于游戏、虚拟交流、教育和内容创作。然而,现有方法往往难以捕捉精细表情,且在不同表情间难以保持身份一致性。本文提出GEN-AFFECT框架,通过将多模态扩散Transformer基于提取的身份-表情表征进行条件化,实现表达丰富且身份一致的虚拟形象生成。该框架在推理时引入一致注意力机制,促进不同表情间的特征共享,从而在生成一系列细粒度表情时维持身份连贯性。实验表明,GEN-AFFECT在表情准确性、身份保留度及跨表情的身份一致性方面均优于现有最先进方法。
原文摘要 · Abstract (English)
Different forms of customized 2D avatars are widely used in gaming applications, virtual communication, education, and content creation. However, existing approaches often fail to capture fine-grained facial expressions and struggle to preserve identity across different expressions. We propose GEN-AFFECT, a novel framework for personalized avatar generation that generates expressive and identity-consistent avatars with a diverse set of facial expressions. Our framework proposes conditioning a multimodal diffusion transformer on an extracted identity-expression representation. This enables identity preservation and representation of a wide range of facial expressions. GEN-AFFECT additionally employs consistent attention at inference for information sharing across the set of generated expressions, enabling the generation process to maintain identity consistency over the array of generated fine-grained expressions. GEN-AFFECT demonstrates superior performance compared to previous state-of-the-art methods on the basis of the accuracy of the generated expressions, the preservation of the identity and the consistency of the target identity across an array of fine-grained facial expressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。