arXiv:2410.12567eess.AScs.SD2024-10

通过逐类微调缓解语音情感识别中的灾难性遗忘问题。

SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning

  • 按情绪类别逐个增量微调模型,避免遗忘旧类
  • 在多个数据集上准确率和F1分数显著优于现有方法
  • 适合多语言、持续学习场景下的语音情感识别

本文提出SeQuiFi,一种用于缓解语音情感识别中灾难性遗忘的新方法。该方法采用顺序类别微调策略,每次仅对一个情绪类别进行增量微调,从而有效保留并增强各类别的识别能力。尽管已有多种先进方法(如正则化、记忆存储、权重平均等)尝试解决此问题,但在多样性和多语言数据集上仍具挑战。大量实验表明,SeQuiFi在CREMA-D、RAVDESS、Emo-DB、MESD和SHEMO等多个基准数据集上,无论准确率还是F1分数均显著优于原始微调及当前最优持续学习技术。

原文摘要 · Abstract (English)

In this work, we introduce SeQuiFi, a novel approach for mitigating catastrophic forgetting (CF) in speech emotion recognition (SER). SeQuiFi adopts a sequential class-finetuning strategy, where the model is fine-tuned incrementally on one emotion class at a time, preserving and enhancing retention for each class. While various state-of-the-art (SOTA) methods, such as regularization-based, memory-based, and weight-averaging techniques, have been proposed to address CF, it still remains a challenge, particularly with diverse and multilingual datasets. Through extensive experiments, we demonstrate that SeQuiFi significantly outperforms both vanilla fine-tuning and SOTA continual learning techniques in terms of accuracy and F1 scores on multiple benchmark SER datasets, including CREMA-D, RAVDESS, Emo-DB, MESD, and SHEMO, covering different languages.

语音情感识别持续学习灾难性遗忘增量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。