构建大规模跨领域语音识别基准,评估模型理解上下文能力。
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
- 设计包含超4万条数据的多领域语音数据集,覆盖30万+命名实体。
- 引入三种评估模式,检验模型利用上下文提升识别准确率的能力。
- 揭示大语言模型在上下文理解上显著优于传统语音模型,仍有提升空间。
自动语音识别(ASR)研究广泛,但现有基准多聚焦声学鲁棒性,对语言能力评估不足。这源于传统ASR模型参数量与训练语料有限,缺乏足够的世界知识,难以准确识别医学、工程等领域的专业术语。近期大语言模型(LLM)及大音频语言模型(LALM)的发展,显著提升了上下文建模与通用智能能力。为此,我们提出ContextASR-Bench:一个涵盖10余个领域、总计达40,000条数据、包含超过300,000个命名实体的大规模基准。每条样本提供音频、转录文本、所属领域及包含的命名实体列表作为上下文信息,并设计三种评估模式,考察模型利用上下文提升识别效果的能力。实验表明,依托强大世界知识和上下文建模能力,LALMs显著优于传统ASR模型,但仍存在明显改进空间。数据集与评测代码已开源。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for accurately recognizing named entities across diverse domains. For instance, drug and treatment names in medicine or specialized technical terms in engineering. Recent breakthroughs in Large Language Models (LLMs) and corresponding Large Audio Language Models (LALMs) have markedly enhanced the visibility of advanced context modeling and general artificial intelligence capabilities. Leveraging LLMs, we envision a unified system capable of robust speech recognition across diverse real-world domains, yet existing benchmarks are inadequate for evaluating this objective. To address this gap, we propose ContextASR-Bench: a comprehensive, large-scale benchmark designed to assess the linguistic competence of ASR systems using corpora that feature numerous named entities across multiple domains. It encompasses up to 40,000 data entries with more than 300,000 named entities across over 10 domains. Beyond the audio and its transcription, each sample provides the domain it belongs to and a list of named entities it contains, which are referred to as the context. Based on this, we introduce three evaluation modes to assess how effectively models can exploit such context to improve ASR accuracy. Extensive evaluation on ContextASR-Bench highlights that LALMs outperform conventional ASR models by a large margin thanks to the strong world knowledge and context modeling of LLMs, yet there remains ample room for further improvement. The dataset and evaluation code have been released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。