用可学习系数融合语音情感与语音识别任务向量,解决模型冲突问题。
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
- 分层自适应融合任务向量,避免梯度干扰
- 在MSP-Podcast上实现更优的情感识别性能
- 适合需融合语言与语调信息的研究者
将自动语音识别(ASR)融入语音情感识别(SER)能提供语言上下文,增强建模效果。然而,传统特征融合存在性能瓶颈,多任务学习常面临优化冲突。尽管任务向量与模型合并已在NLP和计算机视觉中缓解此类问题,其在语音任务中的潜力仍待挖掘。本文提出基于WavLM-Large的自适应分层任务向量合并(AdaLTM)框架。不采用联合优化,而是从领域内微调的ASR与SER模型中提取任务向量,并通过分层可学习系数将其整合到冻结的基模型中。该策略实现了变压器各层间语言与副语言知识的深度感知平衡,避免梯度干扰。在MSP-Podcast数据集上的实验表明,所提方法有效缓解了ASR与SER之间的冲突。
原文摘要 · Abstract (English)
Integrating Automatic Speech Recognition (ASR) into Speech Emotion Recognition (SER) enhances modeling by providing linguistic context. However, conventional feature fusion faces performance bottlenecks, and multi-task learning often suffers from optimization conflicts. While task vectors and model merging have addressed such conflicts in NLP and CV, their potential in speech tasks remains largely unexplored. In this work, we propose an Adaptive Layer-wise Task Vector Merging (AdaLTM) framework based on WavLM-Large. Instead of joint optimization, we extract task vectors from in-domain ASR and SER models fine-tuned on emotion datasets. These vectors are integrated into a frozen base model using layer-wise learnable coefficients. This strategy enables depth-aware balancing of linguistic and paralinguistic knowledge across transformer layers without gradient interference. Experiments on the MSP-Podcast demonstrate that the proposed approach effectively mitigates conflicts between ASR and SER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。