arXiv:2511.13225cs.CL2025-11中稿 · IJCNLP-AACL 2025

测试视觉语言模型解读语音谱图的能力,发现其表现远低于预期。

Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms

  • 构建4000+个孤立英语词的语音谱图与波形数据集
  • 零样本和微调模型准确率仅略高于随机猜测
  • 说明单纯配对数据无法让模型学会解读语音图像

随着大语言模型及其视觉增强版本的兴起,众多研究探讨了多模态融合任务中的能力。本文评估视觉语言模型(VLMs)作为专业语音学家解读语音谱图与波形的能力。为此,我们构建了一个新数据集,包含4000多个孤立英语词的语音信号及风格一致的谱图与波形图像。通过多项选择任务,测试模型在三个基于音素编辑距离设计的干扰项中识别正确音素或字形转录的能力。结果显示,无论是零样本还是微调模型,准确率均远低于显著性水平,表明仅靠配对样本不足以掌握谱图解读所需的专门知识。

原文摘要 · Abstract (English)

With the rise of Large Language Models (LLMs) and their vision-enabled counterparts (VLMs), numerous works have investigated their capabilities in tasks that fuse the modalities of vision and language. In this work, we benchmark the extent to which VLMs are able to act as highly-trained phoneticians, interpreting spectrograms and waveforms of speech. To do this, we synthesise a novel dataset containing 4k+ English words spoken in isolation alongside stylistically consistent spectrogram and waveform figures. We test the ability of VLMs to understand these representations of speech through a multiple-choice task whereby models must predict the correct phonemic or graphemic transcription of a spoken word when presented amongst 3 distractor transcriptions that have been selected based on their phonemic edit distance to the ground truth. We observe that both zero-shot and finetuned models rarely perform above chance, demonstrating the requirement for specific parametric knowledge of how to interpret such figures, rather than paired samples alone.

视觉语言模型语音分析多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。