arXiv:2504.07945cs.CVcs.AI2025-04被引 3

用扩散模型生成细节丰富的卡通头像,支持135种表情且不泄露真实身份。

GenEAva: Generating Cartoon Avatars with Fine-Grained Facial Expressions from Realistic Diffusion-based Faces

  • 微调扩散模型生成逼真表情,再通过风格化转为卡通形象。
  • 构建首个包含13230张卡通头像的数据集,覆盖135种精细表情。
  • 生成结果无真实身份记忆,适合隐私敏感的社交与游戏应用。

卡通头像广泛应用于社交媒体、在线教学和游戏中。然而,现有数据集与生成方法难以呈现具有精细面部表情的高表达力头像,且常基于真实人物,引发隐私问题。为此,我们提出新框架GenEAva,可生成高质量、表情细腻的卡通头像。该方法微调先进的文生图扩散模型以合成高度细节化且富有表现力的面部表情,随后利用风格化模型将这些真实人脸转化为卡通形象,同时保留身份与表情特征。基于此框架,我们构建首个表达性卡通头像数据集GenEAva 1.0,专门捕捉135种精细面部表情,包含13,230张均衡分布于性别、种族与年龄范围内的表达性卡通头像。实验表明,微调后的模型生成的表情比当前最先进的文生图模型SDXL更具表现力;同时验证了生成头像未包含微调数据中的记忆化真实身份。所提框架与数据集为未来卡通头像生成研究提供了多样且富有表现力的基准。

原文摘要 · Abstract (English)

Cartoon avatars have been widely used in various applications, including social media, online tutoring, and gaming. However, existing cartoon avatar datasets and generation methods struggle to present highly expressive avatars with fine-grained facial expressions and are often inspired from real-world identities, raising privacy concerns. To address these challenges, we propose a novel framework, GenEAva, for generating high-quality cartoon avatars with fine-grained facial expressions. Our approach fine-tunes a state-of-the-art text-to-image diffusion model to synthesize highly detailed and expressive facial expressions. We then incorporate a stylization model that transforms these realistic faces into cartoon avatars while preserving both identity and expression. Leveraging this framework, we introduce the first expressive cartoon avatar dataset, GenEAva 1.0, specifically designed to capture 135 fine-grained facial expressions, featuring 13,230 expressive cartoon avatars with a balanced distribution across genders, racial groups, and age ranges. We demonstrate that our fine-tuned model generates more expressive faces than the state-of-the-art text-to-image diffusion model SDXL. We also verify that the cartoon avatars generated by our framework do not include memorized identities from fine-tuning data. The proposed framework and dataset provide a diverse and expressive benchmark for future research in cartoon avatar generation.

卡通头像扩散模型表情生成隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。