arXiv:2508.12255cs.CLeess.AS2025-08被引 2

剖析语音大模型学了什么,推动其在语言理解任务中的应用

What do Speech Foundation Models Learn? Analysis and Applications

  • 用无训练任务和统计工具分析语音模型各层的声学与语言知识
  • 发现端到端模型在语音命名实体识别上优于传统分步方法
  • 开源新任务与数据集,助力语音理解研究发展

语音基础模型(SFMs)旨在为多种语音处理任务提供通用表征。过去五年中,自监督与监督预训练模型不断涌现,在下游任务中表现优异。然而,对这些模型所学习知识的理解仍滞后于模型发展。本文提出一种轻量级分析框架,结合统计工具与无训练任务,研究SFMs各层编码的声学与语言知识,并在多个SFMs与工具间进行对比。研究发现,分析结果直接影响下游任务性能。尽管模型效果最终取决于实际应用表现,但其在需深层理解的语音语言理解(SLU)任务中的有效性仍不明确,主要因缺乏相关数据集。为此,本文贡献了新任务——口语命名实体识别(NER)与实体定位(NEL),并加入语音语言理解评估基准。基于SFMs构建的端到端(E2E)模型在两项任务上超越传统级联(先识别语音再处理文本)方法。进一步评估不同SFMs与适配策略对任务性能的影响。本研究解答了关于SFMs的若干未解问题,提供了分析工具与数据集,推动社区在模型设计与采纳上的科学决策。

原文摘要 · Abstract (English)

Speech foundation models (SFMs) are designed to serve as general-purpose representations for a wide range of speech-processing tasks. The last five years have seen an influx of increasingly successful self-supervised and supervised pre-trained models with impressive performance on various downstream tasks. Although the zoo of SFMs continues to grow, our understanding of the knowledge they acquire lags behind. This thesis presents a lightweight analysis framework using statistical tools and training-free tasks to investigate the acoustic and linguistic knowledge encoded in SFM layers. We conduct a comparative study across multiple SFMs and statistical tools. Our study also shows that the analytical insights have concrete implications for downstream task performance. The effectiveness of an SFM is ultimately determined by its performance on speech applications. Yet it remains unclear whether the benefits extend to spoken language understanding (SLU) tasks that require a deeper understanding than widely studied ones, such as speech recognition. The limited exploration of SLU is primarily due to a lack of relevant datasets. To alleviate that, this thesis contributes tasks, specifically spoken named entity recognition (NER) and named entity localization (NEL), to the Spoken Language Understanding Evaluation benchmark. We develop SFM-based approaches for NER and NEL, and find that end-to-end (E2E) models leveraging SFMs can surpass traditional cascaded (speech recognition followed by a text model) approaches. Further, we evaluate E2E SLU models across SFMs and adaptation strategies to assess the impact on task performance. Collectively, this thesis tackles previously unanswered questions about SFMs, providing tools and datasets to further our understanding and to enable the community to make informed design choices for future model development and adoption.

语音理解基础模型命名实体识别端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。