arXiv:2412.07406cs.CVcs.MM2024-12被引 1

通过自监督学习音频与视觉对应关系,提升音效推荐准确率。

Learning Self-Supervised Audio-Visual Representations for Sound Recommendations

  • 利用注意力机制融合多分辨率音视频特征,建模跨模态对应关系。
  • 在VGG-Sound数据集上,相关性分类准确率提升18%,音效推荐准确率提升10%。
  • 结合对比学习的表示学习方法,显著增强复杂游戏视频场景下的推荐效果。

我们提出一种新颖的自监督方法,从无标签视频中学习音频与视觉表征,基于二者间的对应关系。该方法使用注意力机制,学习不同分辨率下提取的音视频卷积特征的相对重要性,并利用注意力特征编码音视频输入以反映其对应关系。我们评估了模型所学表征在音视频相关性分类及视觉场景音效推荐上的表现。结果表明,相较于基线模型,注意力模型生成的表征在相关性分类准确率上提升18%,在音效推荐准确率上提升10%(基于VGG-Sound公开数据集)。此外,采用跨模态对比学习训练注意力模型所获得的音视频表征,在VGG-Sound及更具挑战性的游戏录像数据集上进一步提升了推荐性能。

原文摘要 · Abstract (English)

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional features extracted at different resolutions from the audio and visual streams and uses the attention features to encode the audio and visual input based on their correspondence. We evaluated the representations learned by the model to classify audio-visual correlation as well as to recommend sound effects for visual scenes. Our results show that the representations generated by the attention model improves the correlation accuracy compared to the baseline, by 18% and the recommendation accuracy by 10% for VGG-Sound, which is a public video dataset. Additionally, audio-visual representations learned by training the attention model with cross-modal contrastive learning further improves the recommendation performance, based on our evaluation using VGG-Sound and a more challenging dataset consisting of gameplay video recordings.

自监督音视频对齐推荐系统对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。