arXiv:2411.15436cs.CV2024-11被引 13

用时间引导的扩散模型生成全程稳定的说话头像。

ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance

  • 先建模帧间时序特征,再通过对比真实视频对齐细节图。
  • 在多个数据集上显著提升外观、3D、表情和时序一致性。
  • 适合需要高稳定性的虚拟主播、数字人应用。

扩散模型在说话头像生成中表现优异,但因单图生成的固有局限与误差累积,仍存在时序、三维或表情不一致问题。本文提出 ConsistentAvatar 框架,实现完全一致且高保真的说话头像生成。方法不直接将多模态条件输入扩散过程,而是先学习相邻帧间的时序表示。具体地,设计了包含高频特征与显著时变轮廓的时敏细节(TSD)图,并通过时序一致扩散模块,将初始结果的 TSD 对齐至视频真值。最终生成依赖对齐后的 TSD、粗略头部法向与情绪提示嵌入的完整头像。实验表明,对齐后的 TSD 作为时序模式约束,有效抑制误差累积,提升各维度一致性。在多个基准上,本方法优于现有最优模型。

原文摘要 · Abstract (English)

Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation and inherent limitation of single-image generation ability. In this paper, we propose ConsistentAvatar, a novel framework for fully consistent and high-fidelity talking avatar generation. Instead of directly employing multi-modal conditions to the diffusion process, our method learns to first model the temporal representation for stability between adjacent frames. Specifically, we propose a Temporally-Sensitive Detail (TSD) map containing high-frequency feature and contours that vary significantly along the time axis. Using a temporal consistent diffusion module, we learn to align TSD of the initial result to that of the video frame ground truth. The final avatar is generated by a fully consistent diffusion module, conditioned on the aligned TSD, rough head normal, and emotion prompt embedding. We find that the aligned TSD, which represents the temporal patterns, constrains the diffusion process to generate temporally stable talking head. Further, its reliable guidance complements the inaccuracy of other conditions, suppressing the accumulated error while improving the consistency on various aspects. Extensive experiments demonstrate that ConsistentAvatar outperforms the state-of-the-art methods on the generated appearance, 3D, expression and temporal consistency. Project page: https://njust-yang.github.io/ConsistentAvatar.github.io/

说话头像扩散模型时序一致数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。