SPEAR统一学习语音与音频表征,性能超越现有模型。
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
- 融合语音与通用音频的教师模型知识,生成统一表征。
- 在SUPERB上15项任务中12项超越WavLM Large。
- 适合需要通用声学表征的语音与音频研究者。
自监督学习(SSL)显著推进了声学表征学习。然而,现有模型大多仅针对语音或音频事件理解优化,导致两领域间存在持续差距。我们提出SPEAR(SPEech and Audio Representations),一种将专注语音的SSL教师与通用音频的SSL教师互补知识提炼至单一统一模型的框架。SPEAR采用多码本向量量化处理连续教师表征,生成捕捉语义与声学信息的细粒度离散标记。为有效整合异构表征,SPEAR在掩码输入下以非对称预训练损失联合预测这些标记。通过新颖的标记混合机制提升复杂声景下的鲁棒性。大量实验表明,SPEAR持续优于现有统一语音与音频模型。在SUPERB基准上建立新SOTA,15项任务中12项超越WavLM Large,HEAR基准上表现具竞争力。SPEAR成为通用语音与音频表征学习的可靠基础。代码与预训练模型将公开。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has significantly advanced acoustic representation learning. However, most existing models are optimised for either speech or audio event understanding, resulting in a persistent gap between these two domains. We address this gap with SPEAR (SPEech and Audio Representations), a self-supervised framework that distils complementary knowledge from a speech-focused SSL teacher and a general-audio SSL teacher into a single unified model. SPEAR applies multi-codebook vector quantisation to continuous teacher representations to produce fine-grained discrete tokens that capture both semantic and acoustic information. To effectively integrate these heterogeneous representations, SPEAR jointly predicts them given a masked input with an asymmetric pre-training loss. We further improve robustness in complex sound scenes through a novel token mixing mechanism. Extensive experiments demonstrate that SPEAR consistently outperforms existing unified speech and audio models. SPEAR establishes a new state-of-the-art on the SUPERB benchmark, surpassing WavLM Large on 12 of 15 tasks, while achieving competitive performance on the HEAR benchmark. These results position SPEAR as a versatile foundation for general-purpose speech and audio representation learning. The code and pre-trained models will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。