arXiv:2409.11729cs.MMcs.CV2024-09被引 1

用物体信息提升音视频表征,让模型更懂具体物品

DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information

  • 在对比掩码自编码器中加入音视频标签预测损失,增强物体感知
  • 在VGGSound和AudioSet20K上,音视频检索召回率提升1.5%和1.2%
  • 无需人工标注,利用现成模型自动获取音视频物体标签

当前音视频表征学习能捕捉粗粒度物体类别(如“动物”和“乐器”),但难以识别细粒度细节,例如“狗”和“长笛”。为此,我们提出DETECLAP,通过在现有对比音视频掩码自编码器中引入音视频标签预测损失,增强其物体意识。为避免高昂的人工标注成本,我们利用先进的语言-音频模型和物体检测器,从音视频输入中自动提取物体标签。在VGGSound和AudioSet20K数据集上评估音视频检索与分类任务,方法在音到视和视到音检索中分别实现recall@10提升1.5%和1.2%,在音视频分类中准确率提升0.6%。

原文摘要 · Abstract (English)

Current audio-visual representation learning can capture rough object categories (e.g., ``animals'' and ``instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like ``dogs'' and ``flutes'' within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification.

音视频融合细粒度识别自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。