arXiv:2601.17085eess.AScs.LG2026-01中稿 · ICASSP 2026被引 1

用多层融合与声学特征增强,让离散语音标记重拾情感识别性能

Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration

  • 通过注意力机制融合多层离散标记,补全信息缺失
  • 引入openSMILE特征,恢复量化丢失的副语言线索
  • 验证多种编码器效果,为语音情感识别提供实用方案

离散语音标记在存储和语言模型集成方面具有显著优势,但在语音情感识别(SER)中的应用受限于量化过程导致的副语言信息丢失。本文系统研究了离散标记在SER中的表现,基于微调的WavLM-Large模型,量化了不同层级配置与k均值聚类粒度下的性能下降程度。为恢复信息损失,提出两项关键策略:(1) 基于注意力的多层融合,以捕获各层级间的互补信息;(2) 集成openSMILE特征,显式重建副语言线索。同时对比主流神经编解码器标记器(SpeechTokenizer、DAC、EnCodec),分析其与声学特征融合时的行为表现。结果表明,通过多层融合与声学特征整合,离散标记可逼近连续表示在SER任务中的性能表现。

原文摘要 · Abstract (English)

Discrete speech tokens offer significant advantages for storage and language model integration, but their application in speech emotion recognition (SER) is limited by paralinguistic information loss during quantization. This paper presents a comprehensive investigation of discrete tokens for SER. Using a fine-tuned WavLM-Large model, we systematically quantify performance degradation across different layer configurations and k-means quantization granularities. To recover the information loss, we propose two key strategies: (1) attention-based multi-layer fusion to recapture complementary information from different layers, and (2) integration of openSMILE features to explicitly reintroduce paralinguistic cues. We also compare mainstream neural codec tokenizers (SpeechTokenizer, DAC, EnCodec) and analyze their behaviors when fused with acoustic features. Our findings demonstrate that through multi-layer fusion and acoustic feature integration, discrete tokens can close the performance gap with continuous representations in SER tasks.

语音情感识别离散标记多层融合特征集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。