统一训练三模态语义对齐,效果优于分步对齐。
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
- 单阶段联合优化文本、音频、视觉表征
- 音频驱动的视觉检索性能提升一倍
- 适合多模态表示学习研究者参考
本文提出一种单阶段训练方法,通过对比学习框架实现文本、音频和视觉三模态的语义对齐。对比学习在多模态对齐中日益重要,可利用大规模无标签数据学习共享表征。现有深度学习方法通常采用两阶段策略,分别对齐视觉-文本与音频-文本模态,但存在数据分布不匹配问题,导致对齐效果不佳。基于AVCaps数据集(为视频片段提供音频、视觉及音视频字幕),本方法使用对比学习联合优化所有模态的表征。实验结果表明,单阶段方法显著优于两阶段方法,在基于音频的视觉检索任务上实现两倍性能提升,凸显了统一多模态表示学习的优势。
原文摘要 · Abstract (English)
This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。