arXiv:2411.18497cs.SDcs.LG2024-11被引 4

用多选学习替代传统方法,高效解决多人语音分离的匹配难题。

Multiple Choice Learning for Efficient Speech Separation with Many Speakers

  • 引入多选学习框架,避免语音分离中的排列歧义问题。
  • 在WSJ0-mix和LibriMix上性能媲美传统方法,计算开销更低。
  • 可拓展至变人数分离或无监督场景,潜力大。

监督式语音分离训练面临排列问题:需确定模型输出与真实分离信号的最佳对应关系。传统方法采用置换不变训练(PIT)解决此歧义。本文提出采用多选学习(MCL)框架,原用于处理模糊任务。实验表明,在主流的WSJ0-mix和LibriMix数据集上,MCL性能与PIT相当,同时具备计算优势。该方法为未来研究开辟新方向,可自然扩展至可变人数分离或无监督语音分离场景。

原文摘要 · Abstract (English)

Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily solved using Permutation Invariant Training (PIT). In this article, we instead consider using the Multiple Choice Learning (MCL) framework, which was originally introduced to tackle ambiguous tasks. We demonstrate experimentally on the popular WSJ0-mix and LibriMix benchmarks that MCL matches the performances of PIT, while being computationally advantageous. This opens the door to a promising research direction, as MCL can be naturally extended to handle a variable number of speakers, or to tackle speech separation in the unsupervised setting.

语音分离多选学习深度学习音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。