HarmoniFuse让语音模型同时高效处理语音识别与情感识别
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
- 按任务需求选择性融合语音特征,避免干扰
- 在有限数据下同时提升语音识别与情感识别效果
- 适合需要多任务统一建模的语音应用开发
近年来大语言模型的发展推动了统一语音语言模型(SLMs)的出现,可在共享架构中支持多种语音任务。然而,自动语音识别(ASR)和语音情感识别(SER)依赖不同信息:前者侧重语言内容,后者需融合语言与副语言线索。现有多任务SLM通常采用简单参数共享或提示条件化,未显式建模任务间信息组成差异,导致任务干扰与性能下降,尤其在数据受限情况下。为此,我们提出HarmoniFuse——一种组件选择性与提示自适应的多任务语音建模框架。该框架通过门控语音编码器提取任务特定声学特征,并利用提示自适应动态融合模块根据任务特性聚合Transformer层。此外,批次交错训练策略可分别使用ASR与SER数据集,无需联合标注。实验表明,HarmoniFuse在两种任务上均取得提升,为真实数据约束下的多任务语音理解提供了可扩展且鲁棒的解决方案。
原文摘要 · Abstract (English)
Recent advances in large language models have facilitated the development of unified speech language models (SLMs) capable of supporting multiple speech tasks within a shared architecture. However, tasks such as automatic speech recognition (ASR) and speech emotion recognition (SER) rely on distinct types of information: ASR primarily depends on linguistic content, whereas SER requires the integration of both linguistic and paralinguistic cues. Existing multitask SLMs typically adopt naive parameter sharing or prompt-based conditioning without explicitly modeling the differences in information composition required by each task. Such designs risk task interference and performance degradation, especially under limited data conditions. To address these limitations, we propose HarmoniFuse, a component-selective and prompt-adaptive framework for multi-task speech language modeling. HarmoniFuse is designed to harmonize heterogeneous task demands by selecting and fusing task-relevant components of speech representations. Specifically, it integrates a gated speech encoder to extract task-specific acoustic features and a prompt-adaptive dynamic fusion module to aggregate transformer layers based on task characteristics. In addition, a batch-interleaved training strategy enables leveraging separate ASR and SER datasets without requiring joint annotation. Experimental results demonstrate that HarmoniFuse improves both ASR and SER performance, offering a scalable and robust solution for multitask speech understanding under realistic data constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。