arXiv:2607.25289cs.LG2026-07

通过自适应加权与关系对齐,提升轻量语音情感识别性能

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

论文配图:AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
图 1 · 摘自论文原文
  • 根据教师一致性动态分配权重,增强可靠教师贡献
  • 对齐师生间样本关系结构,在多个数据集上超越基线
  • 适合资源受限设备上的实时语音情感识别应用

在设备端进行语音情感识别(SER)对实时应用至关重要,但表现优异的大型自监督模型在边缘设备上过于昂贵。多教师知识蒸馏可将大模型压缩为轻量学生模型,但仍面临两个挑战:教师可靠性随批次变化,且逐标签蒸馏忽略了样本间的相对关系结构。为此,我们提出自适应多教师关系蒸馏(AMRD)。针对每个教师的输出相似性矩阵,使用一类SVM为每批数据分配权重,优先选择更一致的教师。引入关系蒸馏损失,对齐教师与学生之间的相似性矩阵,捕捉传统逐标签匹配所忽略的结构信息。在IEMOCAP和CREMA-D数据集上,针对四种学生架构,AMRD在多数设置下优于单教师蒸馏基线,消融实验验证了两组件均带来互补增益。

原文摘要 · Abstract (English)

On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress them into a lightweight student, but two challenges remain: teacher reliability varies across batches, and logit-level distillation ignores inter-sample relational structure. We propose Adaptive Multi-teacher Relational Distillation (AMRD) to address both. A one-class SVM on each teacher's logit similarity matrix assigns per-batch weights favoring more coherent teachers. A relational distillation loss aligns teacher and student similarity matrices, capturing structure that logit matching misses. On IEMOCAP and CREMA-D datasets across four student architectures, AMRD outperforms single-teacher distillation baselines in most settings, and ablations confirm both components yield complementary gains.

语音情感识别知识蒸馏轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。