arXiv:2507.11967cs.CVeess.AS2025-07被引 2

用自动生成的音视频文本三元组,让模型更懂声音、画面和语言的关联。

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos

  • 用预训练文本编码器增强音视频掩码自编码器,跨模态对齐。
  • 自动从无标注视频生成高质量三元组,提升检索和分类性能。
  • 无需人工标注,适合大规模音视频理解任务研究者使用。

本文提出语言引导的对比音视频掩码自编码器(LG-CAV-MAE),以改进音视频表示学习。该模型将预训练文本编码器融入对比音视频掩码自编码器,实现音频、视觉与文本模态间的联合学习。为训练该模型,我们提出一种自动方法,从无标注视频中生成音视频文本三元组:先用图像字幕模型生成帧级描述,再通过基于CLAP的过滤机制确保音频与字幕强对齐。该方法无需人工标注即可生成高质量三元组。我们在音视频检索与分类任务上评估了该方法,结果表明其显著优于现有方法,在检索任务中召回率@10最高提升5.6%,在分类任务中提升3.2%。

原文摘要 · Abstract (English)

In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrained text encoder into contrastive audio-visual masked autoencoders, enabling the model to learn across audio, visual and text modalities. To train LG-CAV-MAE, we introduce an automatic method to generate audio-visual-text triplets from unlabeled videos. We first generate frame-level captions using an image captioning model and then apply CLAP-based filtering to ensure strong alignment between audio and captions. This approach yields high-quality audio-visual-text triplets without requiring manual annotations. We evaluate LG-CAV-MAE on audio-visual retrieval tasks, as well as an audio-visual classification task. Our method significantly outperforms existing approaches, achieving up to a 5.6% improvement in recall@10 for retrieval tasks and a 3.2% improvement for the classification task.

多模态自监督音视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。