arXiv:2508.20914cs.SDcs.LG2025-08被引 3

无标签学习双耳语音空间特征,提升嘈杂环境下的声源定位性能。

Learning Robust Spatial Representations from Binaural Audio through Feature Distillation

  • 通过特征蒸馏预训练,从双耳语音中无监督提取空间特征。
  • 在混响和噪声环境下,微调后性能优于有监督模型和传统方法。
  • 适合做声源定位任务的初始化,尤其在数据稀缺场景下有效。

近期深度表示学习在多个音频任务中表现优异,但其在多通道音频空间表征学习中的应用仍不充分。本文提出一种基于特征蒸馏的预训练框架,无需标签即可学习鲁棒的双耳语音空间表征。具体而言,从干净双耳语音样本中计算空间特征作为预测目标,再用神经网络从对应增强语音中重建这些特征。预训练完成后,舍弃特征预测器,将编码器权重用于初始化方向到达(DoA)估计模型,并进行微调。实验表明,在混响与噪声环境中,微调后的预训练模型在方向到达估计任务上表现优于全监督模型及经典信号处理方法。

原文摘要 · Abstract (English)

Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is underexplored. We investigate the use of a pretraining stage based on feature distillation to learn a robust spatial representation of binaural speech without the need for data labels. In this framework, spatial features are computed from clean binaural speech samples to form prediction labels. These clean features are then predicted from corresponding augmented speech using a neural network. After pretraining, we throw away the spatial feature predictor and use the learned encoder weights to initialize a DoA estimation model which we fine-tune for DoA estimation. Our experiments demonstrate that the pretrained models show improved performance in noisy and reverberant environments after fine-tuning for direction-of-arrival estimation, when compared to fully supervised models and classic signal processing methods.

空间表征双耳音频无监督学习声源定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。