arXiv:2411.04379eess.AS2024-11被引 2

让语音模型同时学习噪声与语音信息,提升音质评估准确率

A Pre-training Framework that Encodes Noise Information for Speech Quality Assessment

  • 用监督方式编码背景噪声,结合自监督学语音特征
  • 在音质评估任务中表现更优,参数更少
  • 适合需要感知环境噪声的语音应用

自监督学习(SSL)在语音处理领域受到关注,因其能生成对多种下游任务有用的鲁棒表示。现有方法多将语音内容(如音素、说话人身份、情感)嵌入表示,但通过对比和自回归学习使表示对背景噪声不变,限制了其在依赖噪声信息任务中的应用。为此,我们提出一种预训练框架,在监督学习中编码背景噪声信息,同时结合自监督策略嵌入语音内容。实验使用多种编码器,结果表明该框架在感知语音质量评估任务中表现更优,且所需参数更少,优于多个基线方法。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has grown in interest within the speech processing community, since it produces representations that are useful for many downstream tasks. SSL uses global and contextual methods to produce robust representations, where SSL even outperforms supervised models. Most self-supervised approaches, however, are limited to embedding information about, i.e., the phonemes, speaker identity, and emotion, into the extracted representations, where they become invariant to background sounds due to contrastive and auto-regressive learning. This is limiting because many downstream tasks leverage noise information to function accurately. Therefore, we propose a pre-training framework that learns information pertaining to background noise in a supervised manner, while jointly embedding speech information using a self-supervised strategy. We experiment with multiple encoders and show that our framework is useful for perceptual speech quality estimation, which relies on background cues. Our results show that the proposed approach improves performance with fewer parameters, in comparison to multiple baselines.

语音质量自监督学习噪声建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。