用少量语音数据个性化微调模型,显著提升帕金森患者语音增强效果。
Variational Autoencoder for Personalized Pathological Speech Enhancement
- 基于预训练变分自编码器,仅用几秒清洁语音微调个性化模型。
- 个性化模型使正常与帕金森患者语音增强效果接近一致。
- 适合需要适配个体病理语音的临床语音增强场景。
尽管跨说话人泛化性对语音增强(SE)模型的广泛应用至关重要,但该问题仍缺乏充分研究。本文探究混合变分自编码器(VAE)-非负矩阵分解(NMF)模型在帕金森病患者语音增强中的表现。结果表明,基于大规模正常人数据集训练的VAE模型在病理语音上表现不佳;即使通过病理语音微调,仍存在性能差距。为此,我们提出仅用每位说话者几秒干净语音微调预训练模型,构建个性化增强模型。实验显示,个性化模型显著提升所有说话者的性能,使正常与帕金森患者的表现趋于一致。
原文摘要 · Abstract (English)
The generalizability of speech enhancement (SE) models across speaker conditions remains largely unexplored, despite its critical importance for broader applicability. This paper investigates the performance of the hybrid variational autoencoder (VAE)-non-negative matrix factorization (NMF) model for SE, focusing primarily on its generalizability to pathological speakers with Parkinson's disease. We show that VAE models trained on large neurotypical datasets perform poorly on pathological speech. While fine-tuning these pre-trained models with pathological speech improves performance, a performance gap remains between neurotypical and pathological speakers. To address this gap, we propose using personalized SE models derived from fine-tuning pre-trained models with only a few seconds of clean data from each speaker. Our results demonstrate that personalized models considerably enhance performance for all speakers, achieving comparable results for both neurotypical and pathological speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。