提升语音增强模型对病理性语音的适应能力
Generalizability of Predictive and Generative Speech Enhancement Models to Pathological Speakers
- 用病理性语音数据训练或微调模型
- 多说话人微调效果最佳,单人定制受限于数据量
- 为听障人群语音处理提供可行方案
当前最先进的语音增强(SE)模型在典型语音上表现优异,但在病理性语音上的性能显著下降。本文研究了针对预测型和生成型SE模型的改进策略:一、使用病理性语音数据从头训练;二、在典型语音预训练模型基础上,用多个病理性说话人数据进行微调;三、仅用单个病理性说话人的数据进行个性化适配。结果显示,尽管病理性语音数据集规模有限,模型仍可成功训练或微调。多个病理性说话人数据微调带来最大性能提升,而单人个性化适配效果较差,可能因每位说话人可用数据量过少。这些发现揭示了提升病理性语音增强性能所面临的挑战与可行路径。
原文摘要 · Abstract (English)
State of the art speech enhancement (SE) models achieve strong performance on neurotypical speech, but their effectiveness is substantially reduced for pathological speech. In this paper, we investigate strategies to address this gap for both predictive and generative SE models, including i) training models from scratch using pathological data, ii) finetuning models pretrained on neurotypical speech with additional data from pathological speakers, and iii) speaker specific personalization using only data from the individual pathological test speaker. Our results show that, despite the limited size of pathological speech datasets, SE models can be successfully trained or finetuned on such data. Finetuning models with data from several pathological speakers yields the largest performance improvements, while speaker specific personalization is less effective, likely due to the small amount of data available per speaker. These findings highlight the challenges and potential strategies for improving SE performance for pathological speakers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。