让大模型学会识别语气、情绪等语音隐含信息
Resurfacing Paralinguistic Awareness in Large Audio Language Models
- 通过分层分析定位语音中的情感与语境特征层
- 新训练方法让模型在语气识别任务上超越全层微调
- 适合需要理解说话人情绪的语音交互场景
大型音频语言模型(LALMs)拓展了人机交互的语音模态,其潜力源于语音中隐含的副语言线索,能反映用户上下文。然而,当前以内容为中心的范式使LALMs通常忽略这些副语言线索,仅基于查询内容响应。为此,本文提出五种分层分析方法,联合识别副语言特征层与语义理解层,并据此设计一种增强副语言感知的微调协议(PE-FT),包括选择性层微调和辅助双层级分类头。实验表明,该协议能高效、有效地恢复模型对副语言信息的感知能力,甚至优于全层微调策略。
原文摘要 · Abstract (English)
Large Audio Language Models (LALMs) have expanded the interaction with human to speech modality, which introduces great interactive potential, due to the paralinguistic cues implicitly indicating the user context. However, building on the current content-centred paradigm, LALMs usually neglect such paralinguistic cues and respond solely based on query content. In this work, to resurface the paralinguistic awareness in LALMs, we introduce five diverse layer-wise analyses to jointly identify paralinguistic layers and semantic understanding layers. Based on these insights, we propose a paralinguistic-enhanced fine-tuning (PE-FT) protocol accordingly to equip LALMs with paralinguistic-aware capabilities, including (1) selective-layer fine-tuning, and (2) an auxiliary dual-level classification head. Our experiments demonstrate that PE-FT protocol efficiently and effectively resurfaces the paralinguistic awareness, even surpassing the performance of the all-layer fine-tuning strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。