一个语音模型同时学多种说话人属性,提升多语言语音检索效果。
Learning Multiple Utterance-Level Attribute Representations with a Unified Speech Encoder

- 用统一框架让模型生成多种说话人级语义表示
- 在多语言语音检索和说话人识别任务上表现更优
- 适合需要多属性理解的语音应用开发
基于自监督学习训练的语音基础模型生成通用语音表征,可支持多种语音处理任务。通过有监督微调,这些模型在特定下游任务中表现优异。近期的后训练方法(如 SAMU-XSLR 和 SONAR)将语音表征与说话人级语义表征对齐,推动了多模态(语音-文本)和多语言应用的发展。尽管语音基础模型通常在声学帧级别学习上下文嵌入,这些方法则在说话人级别学习表征。本文拓展此范式,提出一种统一的后训练框架,使单个语音基础模型能够生成多种类型的说话人级表征。我们通过联合学习语义与说话人表征,在多语言语音检索和说话人识别任务上验证了该方法的有效性。
原文摘要 · Abstract (English)
Speech foundation models trained with self-supervised learning produce generic speech representations that support a wide range of speech processing tasks. When further adapted with supervised learning, these models can achieve strong performance on specific downstream tasks. Recent post-training approaches, such as SAMU-XSLR and SONAR, align speech representations with utterance-level semantic representations, enabling effective multimodal (speech-text) and multilingual applications. While speech foundation models typically learn contextual embeddings at the acoustic frame level, these methods learn representations at the utterance level. In this work, we extend this paradigm to arbitrary utterance-level attributes and propose a unified post-training framework that enables a single speech foundation model to generate multiple types of utterance-level representations. We demonstrate the effectiveness of this approach by jointly learning semantic and speaker representations and evaluating them on multilingual speech retrieval and speaker recognition tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。