arXiv:2506.00350cs.SDeess.AS2025-06中稿 · Interspeech 2025被引 7

用扩散模型提升失语语音重建质量,更清晰且保留说话人特征。

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model

  • 基于潜空间扩散模型,分三步重建语音内容与说话人身份。
  • 在UASpeech数据集上,语音可懂度和说话人相似性显著提升。
  • 适合语音康复、残障人士语音辅助系统研究者参考。

失语语音重建(DSR)旨在将失语语音转化为清晰可懂的语音,同时保持说话人特征。尽管已有进展,现有方法仍常面临语音可懂度低、说话人相似性差的问题。本文提出一种基于扩散模型的新型DSR系统,利用潜空间扩散模型提升重建语音质量。模型包含三个组件:(i) 语音内容编码器,通过预训练自监督学习(SSL)语音基础模型恢复音素嵌入;(ii) 说话人身份编码器,采用上下文学习机制实现说话人感知的身份保留;(iii) 基于扩散的语音生成器,根据恢复的音素嵌入和保留的说话人身份重建语音。在广泛使用的UASpeech数据集上的评估表明,所提模型在语音可懂度和说话人相似性方面均有显著提升。

原文摘要 · Abstract (English)

Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, existing methods often struggle with low speech intelligibility and poor speaker similarity. In this study, we introduce a novel diffusion-based DSR system that leverages a latent diffusion model to enhance the quality of speech reconstruction. Our model comprises: (i) a speech content encoder for phoneme embedding restoration via pre-trained self-supervised learning (SSL) speech foundation models; (ii) a speaker identity encoder for speaker-aware identity preservation by in-context learning mechanism; (iii) a diffusion-based speech generator to reconstruct the speech based on the restored phoneme embedding and preserved speaker identity. Through evaluations on the widely-used UASpeech corpus, our proposed model shows notable enhancements in speech intelligibility and speaker similarity.

语音重建扩散模型失语语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。