arXiv:2409.09785cs.CLcs.AI2024-09被引 16

用大模型改进语音识别后的文本纠错、说话人标注和情感识别。

Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition

  • 用预训练大模型对语音识别结果进行后处理纠错
  • 在三个任务上验证了大模型提升语音理解能力的潜力
  • 适合研究语音+大模型融合或智能语音助手的开发者

随着生成式AI技术的发展,一个关键问题是:大语言模型(LLM)如何利用冻结的预训练自动语音识别(ASR)模型输出的文本解码结果来增强声学建模任务。为探索语言模型在语音处理中的新能力,我们提出了生成式语音转录纠错(GenSEC)挑战。该挑战包含三个后ASR语言建模任务:(i)后ASR转录纠错,(ii)说话人标注,(iii)情感识别。这些任务旨在模拟未来基于大模型的智能体处理语音界面的情景,同时通过使用开源预训练语言模型或基于代理的API,保持对广大研究者的可及性。我们还分享了基线评估的见解以及未来评测设计的经验教训。

原文摘要 · Abstract (English)

Given recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations.

语音识别大模型纠错多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。