arXiv:2506.16127cs.SDcs.AI2025-06中稿 · Interspeech 2025被引 2

用流匹配模型提升失语症语音可懂度,速度更快更清晰。

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching

  • 采用条件流匹配与扩散变换器,非自回归直接转换失语语音
  • 离散声学单元使可懂度显著提升,收敛速度比梅尔频谱快
  • 基于WavLM提取特征,有效减少说话人差异影响

失语症是一种神经性疾病,严重损害语音可懂度,常导致患者无法有效沟通。因此亟需发展可靠的失语症语音转正常语音技术。本文研究自监督学习(SSL)特征及其量化表示作为梅尔频谱的替代方案在语音生成中的应用,并探索通过从WavLM提取特征,在单说话人语音下生成清晰语音以缓解说话人差异。为此,我们提出一种完全非自回归方法,利用条件流匹配(CFM)与扩散变换器,学习失语症语音到清晰语音的直接映射。结果表明,离散声学单元能显著提升可懂度,且收敛速度优于传统梅尔频谱方法。

原文摘要 · Abstract (English)

Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates the development of robust dysarthric-to-regular speech conversion techniques. In this work, we investigate the utility and limitations of self-supervised learning (SSL) features and their quantized representations as an alternative to mel-spectrograms for speech generation. Additionally, we explore methods to mitigate speaker variability by generating clean speech in a single-speaker voice using features extracted from WavLM. To this end, we propose a fully non-autoregressive approach that leverages Conditional Flow Matching (CFM) with Diffusion Transformers to learn a direct mapping from dysarthric to clean speech. Our findings highlight the effectiveness of discrete acoustic units in improving intelligibility while achieving faster convergence compared to traditional mel-spectrogram-based approaches.

语音合成失语症流匹配非自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。