arXiv:2501.13465cs.SDeess.AS2025-01被引 6

用统一模型同时完成语音增强和神经声码器任务

Neural Vocoders as Speech Enhancers

  • 发现语音增强与声码器在秩退化行为上具有共性,可共享模型架构
  • 单一联合训练模型在两项任务上表现接近独立训练模型
  • 为语音重建提供统一框架,适合研究语音生成与修复的学者

语音增强(SE)与神经声码器传统上被视为独立任务。本文观察到两者在过程秩行为上存在共同特征,由此提出两个关键问题:能否将一个任务的模型用于另一任务的秩退化?是否可用统一模型同时处理两项任务?实验表明,现有语音增强模型可成功用于声码任务;而联合训练的单一模型在两项任务上的表现与分别训练的模型相当。结果表明,语音增强与神经声码器可在更广泛的语音重建框架下统一。代码已开源。

原文摘要 · Abstract (English)

Speech enhancement (SE) and neural vocoding are traditionally viewed as separate tasks. In this work, we observe them under a common thread: the rank behavior of these processes. This observation prompts two key questions: \textit{Can a model designed for one task's rank degradation be adapted for the other?} and \textit{Is it possible to address both tasks using a unified model?} Our empirical findings demonstrate that existing speech enhancement models can be successfully trained to perform vocoding tasks, and a single model, when jointly trained, can effectively handle both tasks with performance comparable to separately trained models. These results suggest that speech enhancement and neural vocoding can be unified under a broader framework of speech restoration. Code: https://github.com/Andong-Li-speech/Neural-Vocoders-as-Speech-Enhancers.

语音增强神经声码器统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。