arXiv:2501.07329cs.SDcs.CL2025-01中稿 · ICASSP 2025被引 1

同时完成语音识别与结构化语义理解,提升听觉语言理解效果。

Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding

  • 基于跨度的端到端框架,联合建模语音识别与语义结构提取。
  • 在中文AISHELL-NER和英文SLURP数据集上均达到当前最优性能。
  • 适合需要实时语音理解的场景,如智能助手、语音交互系统。

语音理解(SLU)是语音领域中的结构预测任务。近年来,许多将SLU视为序列到序列任务的工作取得了显著成果。然而,该方法不适用于同时进行语音识别与理解的场景。本文提出一种联合语音识别与结构学习框架(JSRSL),这是一种基于跨度的端到端SLU模型,能够准确转录语音并同时提取结构化内容。我们在中文数据集AISHELL-NER和英文数据集SLURP上进行了实验,结果表明,所提方法不仅在转录和信息提取能力上优于传统序列到序列方法,还在两个数据集上达到了当前最优表现。

原文摘要 · Abstract (English)

Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved great success. However, This method is not suitable for simultaneous speech recognition and understanding. In this paper, we propose a joint speech recognition and structure learning framework (JSRSL), an end-to-end SLU model based on span, which can accurately transcribe speech and extract structured content simultaneously. We conduct experiments on name entity recognition and intent classification using the Chinese dataset AISHELL-NER and the English dataset SLURP. The results show that our proposed method not only outperforms the traditional sequence-to-sequence method in both transcription and extraction capabilities but also achieves state-of-the-art performance on the two datasets.

语音理解端到端结构学习多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。