arXiv:2410.14997cs.SDcs.AI2024-10中稿 · ICASSP 2025被引 10

通过合成母语发音真值数据,提升非母语者语音口音与发音准确度。

Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS

  • 用母语发音生成理想真值音频,指导模型直接学习口音转换映射。
  • 在保持原说话人身份和语调的前提下,显著改善发音准确性。
  • 适用于需要精准语音矫正的教育或语音助手场景。

以往的口音转换方法主要关注使非母语语音听起来更像母语者,同时保持内容与说话人身份。然而,非母语者常存在发音问题,影响理解。为此,本文提出一种新方法,不仅实现口音转换,还改善发音质量。给定非母语语音及对应文本,系统生成具有母语发音、保留原始时长与韵律的理想真值音频,用于训练模型学习从有口音到标准发音的直接映射。采用端到端 VITS 框架实现高质量波形重建。实验表明,该系统不仅能生成接近母语口音且保留原说话人特征的语音,还能有效提升发音准确性。

原文摘要 · Abstract (English)

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence, we developed a new AC approach that not only focuses on accent conversion but also improves pronunciation of non-native accented speaker. By providing the non-native audio and the corresponding transcript, we generate the ideal ground-truth audio with native-like pronunciation with original duration and prosody. This ground-truth data aids the model in learning a direct mapping between accented and native speech. We utilize the end-to-end VITS framework to achieve high-quality waveform reconstruction for the AC task. As a result, our system not only produces audio that closely resembles native accents and while retaining the original speaker's identity but also improve pronunciation, as demonstrated by evaluation results.

口音转换语音生成发音矫正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。