arXiv:2411.16748cs.CV2024-11被引 1

用记忆库和多模态融合,实现长时长真人说话视频的高质量生成。

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

  • 引入噪声正则化记忆库,缓解长时间生成中的误差累积。
  • 生成视频保持高一致性与流畅性,参数量减少8倍。
  • 适合需要真实生动长视频生成的AI应用开发人员。

长时长说话视频合成在视频质量、人物一致性、时间连贯性和计算效率方面面临持续挑战。随着视频长度增加,视觉退化、人物漂移、时间伪影和误差累积等问题日益严重,严重影响结果的真实感与可靠性。为此,我们提出LetsTalk,一种配备多模态引导和新型记忆库机制的扩散变换器框架,明确维持上下文连续性,实现鲁棒、高质量且高效的长时长说话视频生成。具体而言,LetsTalk引入噪声正则化记忆库,缓解扩展生成过程中的误差累积与采样伪影。为提升效率与时空建模能力,采用深度压缩自编码器与时空感知变压器(线性注意力),实现有效的多模态融合。系统分析三种融合策略,发现深层融合(共生融合)用于人物特征、浅层融合(直接融合)用于音频可实现更优的视觉真实感与精准语音驱动动作,同时保持动作多样性。大量实验表明,LetsTalk在生成质量上达到新SOTA,生成具有增强多样性和生动性的时序一致且逼真的说话视频,并以8倍更少参数量保持卓越效率。

原文摘要 · Abstract (English)

Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and computational efficiency. As video length increases, issues such as visual degradation, portrait drift, temporal artifacts, and error accumulation become increasingly problematic, severely affecting the realism and reliability of the results. To address these challenges, we present LetsTalk, a diffusion transformer framework equipped with multimodal guidance and a novel memory bank mechanism, explicitly maintaining contextual continuity and enabling robust, high-quality, and efficient generation of long-duration talking videos. In particular, LetsTalk introduces a noise-regularized memory bank to alleviate error accumulation and sampling artifacts during extended video generation. To further improve efficiency and spatiotemporal modeling, LetsTalk employs a deep compression autoencoder and a spatiotemporal-aware transformer with linear attention for effective multimodal fusion. We systematically analyze three fusion schemes and show that combining deep (Symbiotic Fusion) for portrait features and shallow (Direct Fusion) for audio achieves superior visual realism and precise speech-driven motion, while preserving diversity of movements. Extensive experiments demonstrate that LetsTalk establishes new state-of-the-art in generation quality, producing temporally coherent and realistic talking videos with enhanced diversity and liveliness, and maintains remarkable efficiency with 8x fewer parameters than previous approaches.

视频生成扩散模型多模态记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。