用语义学习提升原始音频模型鲁棒性,实现低延迟高效音频表征。
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
- 基于联合嵌入架构,通过高层语义学习替代语音单元级表征。
- 在多个下游任务上超越现有时域模型,且计算开销更小。
- 多通道版本对混响和噪声环境有强鲁棒性,适合真实场景应用。
从原始波形学习音频表征克服了基于频谱图方法的局限,如频谱计算延迟长、相位信息丢失。尽管基于原始波形的自监督语音表征学习已取得显著进展,但通用音频表征学习仍未能达到类似效果。本文提出WavJEPA,一种基于波形的联合嵌入预测架构,利用高层语义学习解决语音单元或分词级别表征的不足。实验表明,该方法在多种下游任务中显著优于现有最先进的时域音频基础模型,同时计算资源需求更低。为应对时域模型在嘈杂和混响环境中的性能下降,我们进一步提出WavJEPA-Nat——一个在模拟自然场景下训练的多通道扩展架构。结果表明,WavJEPA-Nat对混响和噪声具有高度鲁棒性。这些成果证明了从原始波形进行通用音频表征学习的可行性与计算效率,展示了面向真实应用的低延迟、高鲁棒性时域音频基础模型潜力。
原文摘要 · Abstract (English)
Learning audio representations from raw waveforms overcomes key limitations of spectrogram-based audio representation learning, such as the long latency of spectrogram computation and the loss of phase information. Yet, while self-supervised speech representation learning from raw waveforms has been remarkably successful, these approaches have not achieved similar feats for general-purpose audio representation learning from waveforms. Here, we propose WavJEPA, a waveform-based version of the Joint-Embedding Predictive Architecture. WavJEPA leverages high-level semantic representation learning to tackle the shortcomings of representation learning at the speech unit or token level. We show that this approach substantially outperforms state-of-the-art time-domain audio foundation models across a wide variety of downstream benchmark tasks, while requiring considerably fewer computational resources. Additionally, to overcome the performance drop that time-domain models typically exhibit in noisy and reverberant real-world acoustic environments, we present WavJEPA-Nat. WavJEPA-Nat is a multi-channel extension of the WavJEPA architecture trained on simulated naturalistic scenes. We find that WavJEPA-Nat is highly robust to reverberation and noise. These results highlight the feasibility and computational efficiency of general-purpose audio representation learning from raw waveforms, showcasing the potential for low-latency, robust time-domain audio foundation models for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。