arXiv:2409.05330cs.CVcs.MM2024-09

用双域音频融合生成更自然的说话人脸关键点

KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation

  • 将音频分为情绪与面部上下文两域,分治学习
  • 基于KAN的融合机制提升关键点生成稳定性
  • 适合虚拟助手、虚拟人等音视频交互场景

音频驱动的说话人脸生成因其广泛应用而备受关注,对教育、医疗、在线对话、虚拟助手和虚拟现实等领域具有重要意义。早期研究仅关注口型变化,导致应用受限;近期方法尝试重建完整面部(含脸姿、颈部、肩部),需依赖关键点生成。然而,如何使关键点稳定且与音频精准对齐仍是难点。本文提出基于KAN的双域融合模型KFusion,将音频分离为情绪信息与面部上下文两个域,通过基于KAN的融合机制实现高质量关键点生成。相比现有模型,该方法在生成效率与对齐精度上表现更优,为未来音频驱动说话人脸生成提供了坚实基础。

原文摘要 · Abstract (English)

Audio-driven talking face generation is a widely researched topic due to its high applicability. Reconstructing a talking face using audio significantly contributes to fields such as education, healthcare, online conversations, virtual assistants, and virtual reality. Early studies often focused solely on changing the mouth movements, which resulted in outcomes with limited practical applications. Recently, researchers have proposed a new approach of constructing the entire face, including face pose, neck, and shoulders. To achieve this, they need to generate through landmarks. However, creating stable landmarks that align well with the audio is a challenge. In this paper, we propose the KFusion of Dual-Domain model, a robust model that generates landmarks from audio. We separate the audio into two distinct domains to learn emotional information and facial context, then use a fusion mechanism based on the KAN model. Our model demonstrates high efficiency compared to recent models. This will lay the groundwork for the development of the audio-driven talking face generation problem in the future.

语音驱动人脸生成关键点预测KAN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。