可控强度的稳定人脸说话视频生成,解决闪烁和失真问题。
ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise Search
- 用光流分离表情与动作,减少视觉闪烁。
- 音频转强度模型实现帧级动态控制,同步更自然。
- 噪声初始化策略提升身份一致性和运动连贯性,适合影视合成场景。
近期视频扩散模型在音频驱动人脸动画方面取得显著进展,但现有方法仍存在闪烁、身份漂移和音画不同步等问题,主要源于外观与运动表征耦合及不稳定的推理策略。本文提出 extbf{ConsistTalk},一种可调控强度且时序一致的说话头生成框架,结合扩散噪声搜索推理。首先,设计 extbf{光流引导的时间模块(OFT)},利用面部光流将运动特征与静态外观解耦,降低视觉闪烁并提升时序一致性。其次,提出通过多模态师生知识蒸馏获得的 extbf{音频转强度(A2I)模型},将音频与面部速度特征转化为帧级强度序列,实现音视频运动联合建模,支持细粒度帧级动态控制,同时保持紧密音画同步。第三,引入 extbf{扩散噪声初始化策略(IC-Init)},在推理阶段对背景一致性和运动连续性施加显式约束,相比当前自回归策略,显著提升身份保留能力并优化运动动态。大量实验表明,ConsistTalk 在减少闪烁、保持身份一致性和生成高保真时序稳定视频方面显著优于已有方法。
原文摘要 · Abstract (English)
Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily stem from entangled appearance-motion representations and unstable inference strategies. In this paper, we introduce \textbf{ConsistTalk}, a novel intensity-controllable and temporally consistent talking head generation framework with diffusion noise search inference. First, we propose \textbf{an optical flow-guided temporal module (OFT)} that decouples motion features from static appearance by leveraging facial optical flow, thereby reducing visual flicker and improving temporal consistency. Second, we present an \textbf{Audio-to-Intensity (A2I) model} obtained through multimodal teacher-student knowledge distillation. By transforming audio and facial velocity features into a frame-wise intensity sequence, the A2I model enables joint modeling of audio and visual motion, resulting in more natural dynamics. This further enables fine-grained, frame-wise control of motion dynamics while maintaining tight audio-visual synchronization. Third, we introduce a \textbf{diffusion noise initialization strategy (IC-Init)}. By enforcing explicit constraints on background coherence and motion continuity during inference-time noise search, we achieve better identity preservation and refine motion dynamics compared to the current autoregressive strategy. Extensive experiments demonstrate that ConsistTalk significantly outperforms prior methods in reducing flicker, preserving identity, and delivering temporally stable, high-fidelity talking head videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。