arXiv:2607.23855cs.SDcs.CV2026-07

让音视频生成更同步,靠的是统一的潜在空间对齐。

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

论文配图:OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
图 1 · 摘自论文原文
  • 音视频联合训练VAE,通过对比学习对齐潜在表示
  • 引入分段级对比损失,实现时序语义精准对应
  • 适合需要高同步性音视频生成的研究者

当前生成模型正从单一音或视频转向音视频联合生成。然而,由于两者结构差异大,实现细粒度跨模态对应仍具挑战。现有方法多采用独立训练音视频VAE,导致潜在空间缺乏跨模态对齐,下游模型需从零学习同步关系。本文提出OmniVAE,一种联合训练的音视频变分自编码器,通过分段级音频-视频对比目标捕捉时序-语义对应,实现两模态潜在表示的精细对齐。同时,它将预训练的模态专用语义编码器特征蒸馏到各自模态中,提升潜在空间的可学习性。大量实验表明,两种机制均显著增强潜在空间表达能力,使下游文本到音视频生成的合成质量更高、跨模态同步更准确。结果强调了构建统一表征对多模态建模的基础价值。

原文摘要 · Abstract (English)

Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1

音视频生成跨模态对齐变分自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。