arXiv:2505.13270cs.SDeess.AS2025-05中稿 · INTERSPEECH 2025被引 10

用任务向量插值融合语音与音乐模型,实现高效统一音频理解。

Distilling a speech and music encoder with task arithmetic

  • 分离训练语音与音乐的教师模型,再通过线性插值生成统一模型。
  • 在语音和音乐基准上性能优于集成蒸馏方法,且可灵活调节领域权重。
  • 适合需要通用音频表示的应用,如音频大模型,训练更简单高效。

尽管自监督学习(SSL)在语音和音乐领域取得进展,现有模型仍分别处理这两个领域,限制了统一音频理解能力。对于需要通用表征的应用(如音频大语言模型),统一模型更具优势。然而,直接训练通用模型计算成本高。知识蒸馏教师集成是自然解决方案,但我们认为解耦语音与音乐模型的蒸馏过程更具灵活性。因此,我们提出学习分离的任务向量,并线性插值生成统一的语音+音乐模型。该策略支持通过可调权重灵活控制领域侧重,且训练更简单。在语音和音乐基准上的实验表明,该方法整体性能优于集成蒸馏。

原文摘要 · Abstract (English)

Despite the progress in self-supervised learning (SSL) for speech and music, existing models treat these domains separately, limiting their capacity for unified audio understanding. A unified model is desirable for applications that require general representations, e.g. audio large language models. Nonetheless, directly training a general model for speech and music is computationally expensive. Knowledge Distillation of teacher ensembles may be a natural solution, but we posit that decoupling the distillation of the speech and music SSL models allows for more flexibility. Thus, we propose to learn distilled task vectors and then linearly interpolate them to form a unified speech+music model. This strategy enables flexible domain emphasis through adjustable weights and is also simpler to train. Experiments on speech and music benchmarks demonstrate that our method yields superior overall performance compared to ensemble distillation.

自监督学习音频理解知识蒸馏任务向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。