arXiv:2505.22141cs.CVcs.AI2025-05

让说话头像可编辑,能随意改发型和表情。

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing

  • 分离图像特征与音频驱动,实现表情、发型等属性灵活控制。
  • 在多个数据集上唇动同步准确率优于或相当主流方法。
  • 适合定制虚拟人、在线教学和品牌客服等需要个性化形象的场景。

近年来,语音驱动的说话头像生成取得了显著进展,在唇动同步和情感表达方面表现优异。然而,这些方法大多忽视了面部属性编辑这一关键任务。该能力对于实现深度个性化及拓展实际应用至关重要,涵盖用户定制数字分身、生动在线教育内容以及品牌专属数字客服等领域。在这些关键场景中,对发型、配饰和细微面部特征等视觉属性的灵活调整,是匹配用户偏好、体现多元品牌特质并适应不同情境需求的基础。本文提出 FaceEditTalker,一个统一框架,可在生成高质量语音同步说话头像的同时实现可控的面部属性操作。该方法包含两个核心组件:图像特征空间编辑模块,提取语义与细节特征,支持对表情、发型、配饰等属性的灵活调控;以及语音驱动视频生成模块,将编辑后的特征与音频引导的面部关键点融合,驱动基于扩散模型的生成器。该设计确保了帧间时间一致性、视觉保真度与身份一致性。在公开数据集上的大量实验表明,本方法在唇动同步精度、视频质量与属性可控性方面达到或超越现有代表性基线方法。项目页面:https://peterfanfan.github.io/FaceEditTalker/

原文摘要 · Abstract (English)

Recent advances in audio-driven talking head generation have achieved impressive results in lip synchronization and emotional expression. However, they largely overlook the crucial task of facial attribute editing. This capability is indispensable for achieving deep personalization and expanding the range of practical applications, including user-tailored digital avatars, engaging online education content, and brand-specific digital customer service. In these key domains, flexible adjustment of visual attributes, such as hairstyle, accessories, and subtle facial features, is essential for aligning with user preferences, reflecting diverse brand identities and adapting to varying contextual demands. In this paper, we present FaceEditTalker, a unified framework that enables controllable facial attribute manipulation while generating high-quality, audio-synchronized talking head videos. Our method consists of two key components: an image feature space editing module, which extracts semantic and detail features and allows flexible control over attributes like expression, hairstyle, and accessories; and an audio-driven video generation module, which fuses these edited features with audio-guided facial landmarks to drive a diffusion-based generator. This design ensures temporal coherence, visual fidelity, and identity preservation across frames. Extensive experiments on public datasets demonstrate that our method achieves comparable or superior performance to representative baseline methods in lip-sync accuracy, video quality, and attribute controllability. Project page: https://peterfanfan.github.io/FaceEditTalker/

说话头像属性编辑扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。