arXiv:2412.07754cs.CVcs.AI2024-12被引 7

只需一张参考图和音频,就能生成高保真可控的说话人脸视频。

PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation

  • 采用双网络架构,分别保持身份一致性和动作连贯性。
  • 支持文本提示控制,实现个性化表情与动作生成。
  • 无需大量参考视频,适合实际应用中的快速定制需求。

音频驱动的说话人脸生成在数字通信中极具挑战性。尽管已有显著进展,但多数方法仅关注音频与口型同步,忽视了视觉质量、可定制性和泛化能力等关键因素。为此,我们提出一种新型可定制的一次性音频驱动人脸生成框架PortraitTalk。该框架基于潜在扩散模型,包含两个核心组件:IdentityNet用于保持生成视频帧间身份特征的一致性,AnimateNet则提升时序连贯性与运动一致性。该方法通过结合音频输入与参考图像,降低了对现有方法中常见的参考风格视频的依赖。其关键创新在于引入文本提示,通过解耦交叉注意力机制显著增强生成视频的创意控制力。通过大量实验及新提出的评估指标,模型在生成效果上优于现有最先进方法,为真实应用场景下可定制的逼真说话人脸生成树立了新标准。

原文摘要 · Abstract (English)

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects such as visual quality, customization, and generalization that are crucial to producing realistic talking faces. To address these limitations, we introduce a novel, customizable one-shot audio-driven talking face generation framework, named PortraitTalk. Our proposed method utilizes a latent diffusion framework consisting of two main components: IdentityNet and AnimateNet. IdentityNet is designed to preserve identity features consistently across the generated video frames, while AnimateNet aims to enhance temporal coherence and motion consistency. This framework also integrates an audio input with the reference images, thereby reducing the reliance on reference-style videos prevalent in existing approaches. A key innovation of PortraitTalk is the incorporation of text prompts through decoupled cross-attention mechanisms, which significantly expands creative control over the generated videos. Through extensive experiments, including a newly developed evaluation metric, our model demonstrates superior performance over the state-of-the-art methods, setting a new standard for the generation of customizable realistic talking faces suitable for real-world applications.

说话人脸音频驱动可控生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。