arXiv:2411.09268cs.CV2024-11被引 6

让人脸视频情绪编辑更精细可控,支持多层级情感调整。

LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space

  • 用面部动作单元定义线性情绪空间,实现可解释的情绪向量变换。
  • 通过跨维度注意力网络,精准引导3D人脸模型的可控形变。
  • 适合需要精细情感控制的影视生成、虚拟主播等场景。

现有单次输入说话头生成模型在粗粒度情绪编辑上已取得进展,但在细粒度且可解释的情绪编辑方面仍显不足。我们提出LES-Talker,一种高可解释性的新型单次输入说话头生成模型,实现跨情绪类型、情绪强度和面部单元的细粒度情绪编辑。基于面部动作单元(Facial Action Units)构建线性情绪空间(LES),将情绪转换建模为向量变换。设计跨维度注意力网络(CDAN),深度挖掘LES表示与3D模型表示之间的关联,通过多维特征与结构关系的联合建模,使LES表示能够指导3D模型的可控形变。针对多模态数据偏差问题,采用专用网络结构与训练策略,提升视觉质量。实验表明,该方法在保持高视觉质量的同时,实现了多层次且可解释的细粒度情绪编辑,优于主流方法。

原文摘要 · Abstract (English)

While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be considered fine-grained, it needs to provide clear definitions and sufficiently detailed differentiation. We present LES-Talker, a novel one-shot talking head generation model with high interpretability, to achieve fine-grained emotion editing across emotion types, emotion levels, and facial units. We propose a Linear Emotion Space (LES) definition based on Facial Action Units to characterize emotion transformations as vector transformations. We design the Cross-Dimension Attention Net (CDAN) to deeply mine the correlation between LES representation and 3D model representation. Through mining multiple relationships across different feature and structure dimensions, we enable LES representation to guide the controllable deformation of 3D model. In order to adapt the multimodal data with deviations to the LES and enhance visual quality, we utilize specialized network design and training strategies. Experiments show that our method provides high visual quality along with multilevel and interpretable fine-grained emotion editing, outperforming mainstream methods.

说话头生成情绪编辑线性空间3D人脸

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。