arXiv:2601.16235cs.SDeess.AS2026-01被引 3

用小模型动态优化语音增强中的说话人嵌入,提升个性化效果。

Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

  • 通过对比知识蒸馏训练15万参数的小型编码器
  • 推理时实时优化说话人嵌入,性能显著提升
  • 轻量设计适合移动端部署,适合语音增强研究者

个性化语音增强(PSE)在从干扰语音中提取已知目标语音方面表现优异。现有系统通常将目标语音的表征嵌入增强模型,该表征由上游模型从目标语音的录入片段中提取。然而,这些预训练模型通常较重,且生成的嵌入无法适应推理时目标语音的变化。本文提出一种实时精炼说话人嵌入的方法,采用一个仅15万参数的小型说话人编码器。首先,引入一种新颖的对比知识蒸馏方法,从复杂嵌入中训练该小型编码器;随后在推理阶段将其集成到增强系统中,实验表明该方法在保持低计算开销的同时,显著提升了PSE性能。

原文摘要 · Abstract (English)

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding's quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-thefly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load.

语音增强嵌入优化轻量化对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。