用语音合成嵌入检测非母语发音的重音,准确率提升超16%。
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
- 从语音合成模型中提取声调、能量、时长嵌入,用于重音定位
- 在非母语语料上,重音检测准确率比基线高16.2%(词级)和6.9%(音节级)
- 适合做语言学习辅助系统或语音分析的研究者参考
自动识别词与音节级别的重音对构建计算机辅助语言学习系统至关重要。已有研究表明,当前最先进的文本到语音(TTS)系统所学得的韵律嵌入,可在合成语音中生成与母语者相近的词与音节重音。为评估此类韵律嵌入在非母语语境下的有效性,本研究对比了来自母语与非母语语音的嵌入,分析了由SOTA TTS模型FastSpeech2提取的时长、能量和音高嵌入。实验在两种条件下进行:仅输入文本,以及同时输入语音与文本。前者直接使用TTS推理模式提取嵌入,后者提出在训练模式下提取。实验基于母语语料Tatoeba和非母语语料ISLE,人工标注了词级重音位置。结果表明,在词级与音节级重音检测中,使用TTS嵌入相比启发式特征与自监督Wav2Vec-2.0表征,最高相对提升分别达到13.7% & 5.9% 和16.2% & 6.9%。
原文摘要 · Abstract (English)
Automatic detection of prominence at the word and syllable-levels is critical for building computer-assisted language learning systems. It has been shown that prosody embeddings learned by the current state-of-the-art (SOTA) text-to-speech (TTS) systems could generate word- and syllable-level prominence in the synthesized speech as natural as in native speech. To understand the effectiveness of prosody embeddings from TTS for prominence detection under nonnative context, a comparative analysis is conducted on the embeddings extracted from native and non-native speech considering the prominence-related embeddings: duration, energy, and pitch from a SOTA TTS named FastSpeech2. These embeddings are extracted under two conditions considering: 1) only text, 2) both speech and text. For the first condition, the embeddings are extracted directly from the TTS inference mode, whereas for the second condition, we propose to extract from the TTS under training mode. Experiments are conducted on native speech corpus: Tatoeba, and non-native speech corpus: ISLE. For experimentation, word-level prominence locations are manually annotated for both corpora. The highest relative improvement on word \& syllable-level prominence detection accuracies with the TTS embeddings are found to be 13.7% & 5.9% and 16.2% & 6.9% compared to those with the heuristic-based features and self-supervised Wav2Vec-2.0 representations, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。