AlignDiT通过多模态对齐生成同步自然的语音,提升音视频一致性与发音清晰度。
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
- 基于DiT架构,用三种策略对齐文本、视频与参考音频表示
- 在多个基准上语音质量、同步性与说话人相似度均达领先水平
- 适用于配音、虚拟人等场景,通用性强且效果稳定
本文研究从文本、视频和参考音频等多种输入模态生成高质量语音的任务,该任务在影视制作、配音及虚拟角色等领域具有广泛应用。尽管已有进展,现有方法仍存在语音可懂度低、音视频不同步、语音不自然及参考说话人特征保留不足等问题。为此,我们提出AlignDiT——一种基于上下文学习能力的多模态对齐扩散变换器,通过三种有效策略对齐多模态表示,并引入新颖的多模态无分类器引导机制,使模型在语音合成过程中自适应平衡各模态信息。大量实验表明,AlignDiT在多个基准测试中显著优于现有方法,在语音质量、音视频同步性和说话人相似度方面表现优异。此外,该模型在视频到语音合成、视觉强制对齐等多样化任务中展现出强泛化能力,持续达到最先进性能。演示页面见 https://mm.kaist.ac.kr/projects/AlignDiT。
原文摘要 · Abstract (English)
In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attention due to its wide range of applications, such as film production, dubbing, and virtual avatars. Despite recent progress, existing methods still suffer from limitations in speech intelligibility, audio-video synchronization, speech naturalness, and voice similarity to the reference speaker. To address these challenges, we propose AlignDiT, a multimodal Aligned Diffusion Transformer that generates accurate, synchronized, and natural-sounding speech from aligned multimodal inputs. Built upon the in-context learning capability of the DiT architecture, AlignDiT explores three effective strategies to align multimodal representations. Furthermore, we introduce a novel multimodal classifier-free guidance mechanism that allows the model to adaptively balance information from each modality during speech synthesis. Extensive experiments demonstrate that AlignDiT significantly outperforms existing methods across multiple benchmarks in terms of quality, synchronization, and speaker similarity. Moreover, AlignDiT exhibits strong generalization capability across various multimodal tasks, such as video-to-speech synthesis and visual forced alignment, consistently achieving state-of-the-art performance. The demo page is available at https://mm.kaist.ac.kr/projects/AlignDiT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。