arXiv:2509.08344eess.AS2025-09中稿 · ASRU 2025被引 4

用少量语音样本实现个性化情感识别,提升模型适应能力。

Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model

  • 通过上下文学习,用少量目标说话人语音样本提取个性特征
  • 在新收集的数据集上,性能超越传统个性化方法
  • 适合资源有限场景下快速适配个体情感识别

本文提出一种基于上下文学习(ICL)的语音情感识别(SER)个性化方法。由于情感表达因人而异,针对特定说话人的适配对提升SER性能至关重要。传统方法依赖预先准备的各类情绪语音样本,但实际中难以覆盖所有情绪标签。为此,本文提出在ICL推理过程中,仅用少量目标说话人的情感语音样本作为条件,即可获取其个性特征。该方法基于元训练的语音-语言模型(扩展自大语言模型),学习如何通过ICL实现个性化SER。在新构建的SER数据集上的实验表明,该方法显著优于传统个性化方法。

原文摘要 · Abstract (English)

This paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person, speaker-specific adaptation is crucial for improving the SER performance. Conventional SER methods have been personalized using emotional utterances of a target speaker, but it is often difficult to prepare utterances corresponding to all emotion labels in advance. Our idea to overcome this difficulty is to obtain speaker characteristics by conditioning a few emotional utterances of the target speaker in ICL-based inference. ICL is a method to perform unseen tasks by conditioning a few input-output examples through inference in large language models (LLMs). We meta-train a speech-language model extended from the LLM to learn how to perform personalized SER via ICL. Experimental results using our newly collected SER dataset demonstrate that the proposed method outperforms conventional methods.

情感识别上下文学习少样本语音建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。