提出JOOCI框架,让语音模型同时深度学习内容与表达信息。
JOOCI: a Framework for Learning Comprehensive Speech Representations
- 设计新方法使语音表示同时优化内容与表达特征
- 在超级基准测试中超越WavLM 26.5%且优于同类模型
- 适合需要兼顾说话人识别与语言理解的语音任务
语音信息可分为内容(如语言)和非内容(如说话人、副语言特征)两类。现有自监督学习方法将模型不同层划分为早期专注非内容、后期专注内容的结构,这种分层方式限制了两类信息对深层表示的利用。为此,本文提出JOOCI——一种联合优化内容与非内容信息的新方法,避免在表示深度上妥协。在SUPERB基准的两个说话人识别任务和两个语言任务上,JOOCI相比WavLM提升26.5%,并优于参数量相近(100M)的其他模型,验证了其有效性。
原文摘要 · Abstract (English)
Information in speech can be categorized into two groups: Content (what is being said, such as linguistics) and Other (how it is expressed such as information about speaker and paralinguistic features). Current self-supervised learning (SSL) methods are shown to divide the model's representational-depth or layers in two, with earlier layers specializing in Other and later layers in Content related tasks. This layer-wise division is inherently sub-optimal, as neither information type can use all layers to build hierarchical representations. To address this, we propose JOOCI, a novel speech representation learning method that does not compromise on the representational-depth for either information type. JOOCI outperforms WavLM by 26.5%, and other models of similar size (100M parameters), when evaluated on two speaker recognition and two language tasks from the SUPERB benchmark, demonstrating its effectiveness in Jointly Optimizing Other and Content Information (JOOCI).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。