arXiv:2409.07437cs.SDcs.CL2024-09被引 28

SALMon评测语音模型对噪声、情绪等声学特征的感知能力

Salmon: A Suite for Acoustic Language Model Evaluation

  • 基于建模方法,通过对比正确与错误样本得分评估模型表现
  • 覆盖噪声、情绪、说话人身份和混响四种声学特性,兼顾一致性与文本匹配度
  • 适用于评估大模型在多维度声学信息上的理解能力,开源可复现

语音语言模型近年来展现出作为通用语音处理系统的重要潜力,不仅能捕捉语音内容,还能建模音频中蕴含的丰富声学信息,如情感、背景噪声等。然而,针对多种声学特性的评估基准仍较为缺乏。为此,我们提出SALMon——一个涵盖背景噪声、情感、说话人身份和房间脉冲响应的新型评估套件。该套件不仅评估模型对特定声学元素的一致性,还衡量其与语音文本的匹配程度。采用基于建模的方法,通过判断模型对正确样本的评分是否高于错误样本,实现快速计算,尤其适用于大型模型。我们在SALMon上评估了多个语音语言模型,揭示了各方法的优势与局限。代码与数据已公开于https://pages.cs.huji.ac.il/adiyoss-lab/salmon/

原文摘要 · Abstract (English)

Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existing in audio signals, beyond spoken content, such as emotion, background noise, etc. Despite this, evaluation benchmarks which evaluate awareness to a wide range of acoustic aspects, are lacking. To help bridge this gap, we introduce SALMon, a novel evaluation suite encompassing background noise, emotion, speaker identity and room impulse response. The proposed benchmarks both evaluate the consistency of the inspected element and how much it matches the spoken text. We follow a modelling based approach, measuring whether a model gives correct samples higher scores than incorrect ones. This approach makes the benchmark fast to compute even for large models. We evaluated several speech language models on SALMon, thus highlighting the strengths and weaknesses of each evaluated method. We make the code and data publicly available at https://pages.cs.huji.ac.il/adiyoss-lab/salmon/ .

语音模型声学评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。