arXiv:2504.02302cs.SDcs.LG2025-04

用自监督预测模式让语音分离模型提前‘看’到未来,实现实时流式处理

Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation

  • 设计因果前端,通过预测模式隐含引入未来信息
  • 自监督预训练使模型在混合语音中捕捉预测性特征
  • 适用于实时语音分离场景,提升流式应用性能

语音分离旨在将多说话人混合语音分解为单人语音流。尽管离线方法可实现分离,但不适用于实时流式应用。因果分离模型仅依赖过去和当前信息,适合实时处理,但因缺乏未来上下文导致性能下降。本文提出一种新型前端,通过预测模式隐含引入未来信息,缓解训练与推理间的不匹配。该预训练前端采用因果卷积编码器与Transformer解码器结构,以自监督方式训练,包含自回归混合预测和上下文知识蒸馏两项创新任务,使模型能直接从混合语音中学习预测模式。该前端作为特征提取器生成高质量预测特征。在合成及真实数据集上的全面评估验证了其有效性。

原文摘要 · Abstract (English)

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a processing paradigm is not suitable for real-time streaming applications. Causal separation models, which rely only on past and present information, offer a promising solution for real-time streaming. However, these models typically suffer from notable performance degradation due to the absence of future context. In this paper, we introduce a novel frontend that is designed to mitigate the mismatch between training and run-time inference by implicitly incorporating future information into causal models through predictive patterns. The pretrained frontend employs a transformer decoder network with a causal convolutional encoder as the backbone and is pretrained in a self-supervised manner with two innovative pretext tasks: autoregressive hybrid prediction and contextual knowledge distillation. These tasks enable the model to capture predictive patterns directly from mixtures in a self-supervised manner. The pretrained frontend subsequently serves as a feature extractor to generate high-quality predictive patterns. Comprehensive evaluations on synthetic and real-world datasets validated the effectiveness of the proposed pretrained frontend.

语音分离自监督因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。