arXiv:2507.02900cs.CVcs.AI2025-07综述被引 6

系统梳理人脸对话生成的多模态方法与挑战,助力数字人等应用落地。

Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions

  • 按2D/3D/NeRF/扩散模型等分类梳理生成技术
  • 指出姿态极端、跨语言合成等核心难点
  • 适合数字人、视频配音等场景的开发者参考

说话头生成(THG)已成为计算机视觉中的变革性技术,可实现人脸与图像、音频、文本或视频输入的同步合成。本文全面综述了相关方法与框架,将技术分为基于2D、基于3D、基于神经辐射场(NeRF)、基于扩散模型、参数驱动及其他方法。文章评估了算法、数据集与评价指标,强调感知真实性和技术效率对数字人、视频配音、超低码率视频会议及在线教育等应用的重要性。研究指出依赖预训练模型、极端姿态处理、多语言合成和时序一致性等挑战。未来方向包括模块化架构、多语言数据集、融合预训练与任务特定层的混合模型,以及创新损失函数。通过整合现有研究并探索新兴趋势,本文旨在为该领域研究人员与实践者提供可操作洞察。完整综述、代码与资源列表详见GitHub:https://github.com/VineetKumarRakesh/thg。

原文摘要 · Abstract (English)

Talking Head Generation (THG) has emerged as a transformative technology in computer vision, enabling the synthesis of realistic human faces synchronized with image, audio, text, or video inputs. This paper provides a comprehensive review of methodologies and frameworks for talking head generation, categorizing approaches into 2D--based, 3D--based, Neural Radiance Fields (NeRF)--based, diffusion--based, parameter-driven techniques and many other techniques. It evaluates algorithms, datasets, and evaluation metrics while highlighting advancements in perceptual realism and technical efficiency critical for applications such as digital avatars, video dubbing, ultra-low bitrate video conferencing, and online education. The study identifies challenges such as reliance on pre--trained models, extreme pose handling, multilingual synthesis, and temporal consistency. Future directions include modular architectures, multilingual datasets, hybrid models blending pre--trained and task-specific layers, and innovative loss functions. By synthesizing existing research and exploring emerging trends, this paper aims to provide actionable insights for researchers and practitioners in the field of talking head generation. For the complete survey, code, and curated resource list, visit our GitHub repository: https://github.com/VineetKumarRakesh/thg.

说话头生成数字人多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。