arXiv:2409.08107cs.CLcs.LG2024-09被引 4

将语音识别与开放实体识别联合建模,提升转录准确性和信息量。

WhisperNER: Unified Open Named Entity and Speech Recognition

  • 基于合成语音数据训练,支持开放域实体识别
  • 在跨领域开放实体任务上超越基线模型
  • 适合需要高精度语音转录与实体提取的场景

将命名实体识别(NER)与自动语音识别(ASR)融合可显著提升转录准确性和信息量。本文提出WhisperNER,一种支持联合语音转录与实体识别的新模型。该模型支持开放类型NER,可在推理阶段识别多样且动态变化的实体。我们基于开放NER研究进展,扩充大规模合成数据集,加入合成语音样本,使WhisperNER在大量含多样化NER标签的数据上进行训练。训练时,模型接收实体标签提示,并优化输出带标注实体的转录结果。为评估性能,我们为常用NER基准生成合成语音,并对现有ASR数据集添加开放型NER标签。实验表明,WhisperNER在跨领域开放型NER和监督微调任务中均优于自然基线。

原文摘要 · Abstract (English)

Integrating named entity recognition (NER) with automatic speech recognition (ASR) can significantly enhance transcription accuracy and informativeness. In this paper, we introduce WhisperNER, a novel model that allows joint speech transcription and entity recognition. WhisperNER supports open-type NER, enabling recognition of diverse and evolving entities at inference. Building on recent advancements in open NER research, we augment a large synthetic dataset with synthetic speech samples. This allows us to train WhisperNER on a large number of examples with diverse NER tags. During training, the model is prompted with NER labels and optimized to output the transcribed utterance along with the corresponding tagged entities. To evaluate WhisperNER, we generate synthetic speech for commonly used NER benchmarks and annotate existing ASR datasets with open NER tags. Our experiments demonstrate that WhisperNER outperforms natural baselines on both out-of-domain open type NER and supervised finetuning.

语音识别实体识别开放域联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。