arXiv:2602.06270cs.CL2026-02中稿 · ICLR被引 2

让大模型通过元音级语调特征听懂语音情感,提升识别与解释力。

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

  • 从元音片段提取音高、能量、时长特征,转为自然语言描述
  • 在多数据集上零样本、跨域、跨语言均超越现有方法
  • 支持可解释的情感分析,适合需要透明推理的场景

语音情感识别面临复杂的多模态挑战,需同时理解语言内容与声学表现,尤其是基频、强度和时间动态等韵律特征。尽管大语言模型(LLMs)在基于文本转录的情感识别中展现出推理潜力,但通常忽略细粒度韵律信息,限制了其效果与可解释性。本文提出 VowelPrompt,一种基于语言学的框架,通过可解释的元音级韵律提示增强基于 LLM 的情感识别。借鉴语音学证据——元音是情感韵律的主要载体,VowelPrompt 从对齐的元音段中提取音高、能量和持续时间特征,并将其转化为自然语言描述,以增强可解释性。该设计使 LLM 能联合推理词汇语义与细粒度韵律变化。此外,采用两阶段适配:监督微调(SFT)后接基于可验证奖励的强化学习(RLVR),通过组相对策略优化(GRPO)提升推理能力,强制结构化输出并改善跨领域与说话人泛化性。在多个基准数据集上的广泛评估表明,VowelPrompt 在零样本、微调、跨域与跨语言条件下均持续优于最先进方法,且能生成同时基于上下文语义与细粒度韵律结构的可解释解释。

原文摘要 · Abstract (English)

Emotion recognition in speech presents a complex multimodal challenge, requiring comprehension of both linguistic content and vocal expressivity, particularly prosodic features such as fundamental frequency, intensity, and temporal dynamics. Although large language models (LLMs) have shown promise in reasoning over textual transcriptions for emotion recognition, they typically neglect fine-grained prosodic information, limiting their effectiveness and interpretability. In this work, we propose VowelPrompt, a linguistically grounded framework that augments LLM-based emotion recognition with interpretable, fine-grained vowel-level prosodic cues. Drawing on phonetic evidence that vowels serve as primary carriers of affective prosody, VowelPrompt extracts pitch-, energy-, and duration-based descriptors from time-aligned vowel segments, and converts these features into natural language descriptions for better interpretability. Such a design enables LLMs to jointly reason over lexical semantics and fine-grained prosodic variation. Moreover, we adopt a two-stage adaptation procedure comprising supervised fine-tuning (SFT) followed by Reinforcement Learning with Verifiable Reward (RLVR), implemented via Group Relative Policy Optimization (GRPO), to enhance reasoning capability, enforce structured output adherence, and improve generalization across domains and speaker variations. Extensive evaluations across diverse benchmark datasets demonstrate that VowelPrompt consistently outperforms state-of-the-art emotion recognition methods under zero-shot, fine-tuned, cross-domain, and cross-linguistic conditions, while enabling the generation of interpretable explanations that are jointly grounded in contextual semantics and fine-grained prosodic structure.

语音情感大模型韵律分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。