统一建模文音视三模态,实现高保真同步生成。
3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
- 三模态联合建模,动态融合文本、音频与视频特征。
- 在多个指标上提升音视频同步与三模态对齐效果。
- 可复用预训练文生视频模型,无需修改主干结构。
文生视频扩散模型虽已实现优异视觉质量,但多数系统生成无声视频,且将音频视为次要任务。现有音视频生成流程多为级联阶段,跨模态误差累积且训练目标分离。近期联合生成方法虽缓解此问题,但常依赖双塔架构与人工设计的跨模态桥接,采用静态单次文本条件,难以复用文生视频主干,并难以捕捉音视频与语言随时间的交互。为此,我们提出3MDiT:一种面向文本驱动音视频同步生成的统一三模态扩散变换器。框架将视频、音频与文本视为共同演化的流:音频分支与文生视频主干同构,三模态全连接块在特征层面融合三者信息,可选的动态文本条件机制随音视频证据演化更新文本表示。该设计支持两种模式:从零训练或正交适配预训练文生视频模型,无需修改其主干。实验表明,该方法在多个量化指标上持续提升音视频同步性与三模态对齐,生成高质量视频与真实音频。
原文摘要 · Abstract (English)
Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the task into cascaded stages, which accumulate errors across modalities and are trained under separate objectives. Recent joint audio-video generators alleviate this issue but often rely on dual-tower architectures with ad-hoc cross-modal bridges and static, single-shot text conditioning, making it difficult to both reuse T2V backbones and to reason about how audio, video and language interact over time. To address these challenges, we propose 3MDiT, a unified tri-modal diffusion transformer for text-driven synchronized audio-video generation. Our framework models video, audio and text as jointly evolving streams: an isomorphic audio branch mirrors a T2V backbone, tri-modal omni-blocks perform feature-level fusion across the three modalities, and an optional dynamic text conditioning mechanism updates the text representation as audio and video evidence co-evolve. The design supports two regimes: training from scratch on audio-video data, and orthogonally adapting a pretrained T2V model without modifying its backbone. Experiments show that our approach generates high-quality videos and realistic audio while consistently improving audio-video synchronization and tri-modal alignment across a range of quantitative metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。