arXiv:2601.20573cs.SD2026-01中稿 · IEEE ICASSP 2026

用生成模型解决语音情绪识别,让分类更高效。

Gen-SER: When the generative model meets speech emotion recognition

  • 将情绪标签映射到连续空间,用正弦编码构建目标分布
  • 通过生成模型实现分布迁移,相似度匹配完成分类
  • 方法可拓展至多类语音理解任务,适合新分类场景

语音情绪识别(SER)在语音理解和生成中至关重要。现有方法多基于分类模型或大语言模型。不同于以往方法,本文提出 Gen-SER,将 SER 重构为分布迁移问题,通过生成模型实现。将离散情绪标签投影至连续空间,利用正弦分层编码获得目标终端分布。采用基于目标匹配的生成模型,高效地将初始分布转换为终端分布。分类结果通过计算生成终端分布与真实终端分布的相似度得到。实验验证了该方法的有效性,证明其可扩展至多种语音理解任务,并展现出在更广泛分类任务中的潜力。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) is crucial in speech understanding and generation. Most approaches are based on either classification models or large language models. Different from previous methods, we propose Gen-SER, a novel approach that reformulates SER as a distribution shift problem via generative models. We propose to project discrete class labels into a continuous space, and obtain the terminal distribution via sinusoidal taxonomy encoding. The target-matching-based generative model is adopted to transform the initial distribution into the terminal distribution efficiently. The classification is achieved by calculating the similarity of the generated terminal distribution and ground truth terminal distribution. The experimental results confirm the efficacy of the proposed method, demonstrating its extensibility to various speech-understanding tasks and suggesting its potential applicability to a broader range of classification tasks.

语音识别生成模型情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。