arXiv:2411.12058cs.SDeess.AS2024-11被引 13

用声谱图让视觉语言模型实现少样本音频分类

Vision Language Models Are Few-Shot Audio Spectrogram Classifiers

  • 通过提示模型识别声谱图,实现少样本音频分类
  • GPT-4o在ESC-10数据集上达59.00%准确率
  • 优于商用音频模型和人类专家的视觉分类表现

我们证明了视觉语言模型(VLMs)能够通过对应的声谱图图像识别音频内容。具体而言,通过提供每类的示例声谱图图像进行提示,使VLM在少样本设置下执行音频分类任务。精心设计声谱图表示并选择优质少样本示例后,GPT-4o在ESC-10环境声音分类数据集上实现了59.00%的交叉验证准确率。此外,我们展示了当前VLM在等效音频分类任务中优于唯一可用的具备音频理解能力的商用音频语言模型Gemini-1.5(59.00% vs. 49.62%),甚至在视觉声谱图分类上略胜人类专家(73.75% vs. 72.50% on first fold)。我们设想两个潜在应用场景:(1) 结合VLM的声谱图与语言理解能力用于音频字幕增强;(2) 将视觉声谱图分类作为挑战性任务用于评估VLM。

原文摘要 · Abstract (English)

We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct VLMs to perform audio classification tasks in a few-shot setting by prompting them to classify a spectrogram image given example spectrogram images of each class. By carefully designing the spectrogram image representation and selecting good few-shot examples, we show that GPT-4o can achieve 59.00% cross-validated accuracy on the ESC-10 environmental sound classification dataset. Moreover, we demonstrate that VLMs currently outperform the only available commercial audio language model with audio understanding capabilities (Gemini-1.5) on the equivalent audio classification task (59.00% vs. 49.62%), and even perform slightly better than human experts on visual spectrogram classification (73.75% vs. 72.50% on first fold). We envision two potential use cases for these findings: (1) combining the spectrogram and language understanding capabilities of VLMs for audio caption augmentation, and (2) posing visual spectrogram classification as a challenge task for VLMs.

音频分类视觉语言模型少样本学习声谱图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。