arXiv:2505.14562cs.SDcs.MM2025-05中稿 · European Signal Pr…被引 6

统一训练三模态语义对齐,效果优于分步对齐。

Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities

  • 单阶段联合优化文本、音频、视觉表征
  • 音频驱动的视觉检索性能提升一倍
  • 适合多模态表示学习研究者参考

本文提出一种单阶段训练方法,通过对比学习框架实现文本、音频和视觉三模态的语义对齐。对比学习在多模态对齐中日益重要,可利用大规模无标签数据学习共享表征。现有深度学习方法通常采用两阶段策略,分别对齐视觉-文本与音频-文本模态,但存在数据分布不匹配问题,导致对齐效果不佳。基于AVCaps数据集(为视频片段提供音频、视觉及音视频字幕),本方法使用对比学习联合优化所有模态的表征。实验结果表明,单阶段方法显著优于两阶段方法,在基于音频的视觉检索任务上实现两倍性能提升,凸显了统一多模态表示学习的优势。

原文摘要 · Abstract (English)

This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive training has gained prominence for multimodal alignment, utilizing large-scale unlabeled data to learn shared representations. Existing deep learning approach for trimodal alignment involves two-stages, that separately align visual-text and audio-text modalities. This approach suffers from mismatched data distributions, resulting in suboptimal alignment. Leveraging the AVCaps dataset, which provides audio, visual and audio-visual captions for video clips, our method jointly optimizes the representation of all the modalities using contrastive training. Our results demonstrate that the single-stage approach outperforms the two-stage method, achieving a two-fold improvement in audio based visual retrieval, highlighting the advantages of unified multimodal representation learning.

多模态对齐对比学习三模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。