arXiv:2509.24004cs.CV2025-09

用一张图加文字,生成能精准表达情绪的3D头像。

SIE3D: Single-Image Expressive 3D Avatar Generation via Semantic Embedding and Perceptual Expression Loss

  • 通过语义嵌入融合图像与文本,实现细粒度表情控制。
  • 引入感知表情损失,使生成表情与文本描述匹配度提升。
  • 仅需消费级显卡,生成效果优于现有方法。

从单张图像生成高保真3D头部形象极具挑战性,因现有方法难以通过文本实现精细、直观的表情控制。本文提出SIE3D框架,可基于单张图像和描述性文本生成具有表现力的3D头像。SIE3D通过新颖的条件化机制,将图像中的身份特征与文本的语义嵌入融合,实现细节级控制。为确保生成表情与文本描述一致,引入创新的感知表情损失函数,该函数利用预训练的表情分类器对生成过程进行正则化,保障表情准确性。大量实验表明,SIE3D在身份保留与表情保真度上显著优于竞争方法,且可在单块消费级GPU上运行。

原文摘要 · Abstract (English)

Generating high-fidelity 3D head avatars from a single image is challenging, as current methods lack fine-grained, intuitive control over expressions via text. This paper proposes SIE3D, a framework that generates expressive 3D avatars from a single image and descriptive text. SIE3D fuses identity features from the image with semantic embedding from text through a novel conditioning scheme, enabling detailed control. To ensure generated expressions accurately match the text, it introduces an innovative perceptual expression loss function. This loss uses a pre-trained expression classifier to regularize the generation process, guaranteeing expression accuracy. Extensive experiments show SIE3D significantly improves controllability and realism, outperforming competitive methods in identity preservation and expression fidelity on a single consumer-grade GPU. Project page: https://huang-zhiqi.github.io/SIE3D/

3D生成表情控制单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。