arXiv:2409.09282cs.LGcs.MM2024-09被引 4

通过联合模内与模间对比学习,提升多模态分类性能。

Turbo your multi-modal classification with contrastive learning

  • 同一数据经两次前向传播,生成双表示以构建对比目标。
  • 在语音情感识别任务上达到当前最优性能。
  • 适合需要多模态表征增强的音频文本分类研究者。

对比学习已成为多模态表征学习中最具影响力的范式之一。然而,以往多模态工作主要关注跨模态理解,忽略了模内对比学习,限制了单个模态的表征能力。本文提出一种新颖的对比学习策略——Turbo,通过联合模内与跨模态对比学习来促进多模态理解。具体而言,将多模态数据对通过前向传播两次,采用不同的隐藏层丢弃掩码,得到每个模态的两种不同表示。利用这些表示,构建多个模内和跨模态对比目标进行训练。最终,将自监督的Turbo与有监督的多模态分类相结合,在两个音文分类任务上验证其有效性,实现在语音情感识别基准数据集上的最先进性能。

原文摘要 · Abstract (English)

Contrastive learning has become one of the most impressive approaches for multi-modal representation learning. However, previous multi-modal works mainly focused on cross-modal understanding, ignoring in-modal contrastive learning, which limits the representation of each modality. In this paper, we propose a novel contrastive learning strategy, called $Turbo$, to promote multi-modal understanding by joint in-modal and cross-modal contrastive learning. Specifically, multi-modal data pairs are sent through the forward pass twice with different hidden dropout masks to get two different representations for each modality. With these representations, we obtain multiple in-modal and cross-modal contrastive objectives for training. Finally, we combine the self-supervised Turbo with the supervised multi-modal classification and demonstrate its effectiveness on two audio-text classification tasks, where the state-of-the-art performance is achieved on a speech emotion recognition benchmark dataset.

多模态对比学习语音情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。