让语音识别同时理解语调重点,提升对话语音的语义解析能力
Prominence-aware automatic speech recognition for conversational speech
- 用wav2vec2微调模型检测词级语调突出度
- 在正确识别语句时,突出度检测准确率达85.53%
- 为语言研究和智能对话系统提供语调感知的新范式
本文研究了结合语调突出度检测与语音识别的对话式奥地利德语自动语音识别。首先通过微调wav2vec2模型,开发出词级语调突出度分类器,并用于大规模语料库的韵律突出度自动标注。基于这些标注,训练了能同时转录词语及其突出程度的新颖突出度感知语音识别系统。该系统在识别正确语句时,突出度检测准确率达到85.53%,且整体识别性能与基线系统相当。结果表明,基于Transformer的模型能有效编码韵律信息,为韵律增强型语音识别提供了新贡献,具有语言学研究及韵律驱动对话系统的应用潜力。
原文摘要 · Abstract (English)
This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-tuning wav2vec2 models to classify word-level prominence. The detector was then used to automatically annotate prosodic prominence in a large corpus. Based on those annotations, we trained novel prominence-aware ASR systems that simultaneously transcribe words and their prominence levels. The integration of prominence information did not change performance compared to our baseline ASR system, while reaching a prominence detection accuracy of 85.53% for utterances where the recognized word sequence was correct. This paper shows that transformer-based models can effectively encode prosodic information and represents a novel contribution to prosody-enhanced ASR, with potential applications for linguistic research and prosody-informed dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。