arXiv:2601.16240eess.AScs.LG2026-01中稿 · 2026 IEEE Internat…被引 1

解决语音情绪识别在真实场景中的适应难题。

Test-Time Adaptation for Speech Emotion Recognition

  • 引入无反向传播的测试时自适应方法,仅用无标签目标数据
  • 11种方法对比显示,无反向传播方法效果最佳
  • 情绪表达模糊性导致传统方法失效,适合实际部署场景

语音情绪识别(SER)系统在实际应用中因领域偏移而表现脆弱,如说话人差异、表演式与自然情绪的区别以及跨语料库变化。尽管领域自适应和微调被广泛研究,但它们通常需要源数据或带标签的目标数据,而这在SER中往往不可用或涉及隐私问题。测试时自适应(TTA)通过仅使用无标签目标数据在推理时调整模型,弥补这一空白。然而,现有TTA方法主要针对图像分类和语音识别设计,其在SER独特领域偏移下的有效性尚未被系统评估。本文首次对11种TTA方法在三个典型SER任务上进行系统比较。结果表明,无反向传播的TTA方法最为有效;而熵最小化和伪标签法普遍失败,因其假设存在单一且确定的真实标签,与情绪表达的固有模糊性相悖。此外,没有单一方法在所有场景下均最优,其效果高度依赖于分布偏移类型和具体任务。

原文摘要 · Abstract (English)

The practical utility of Speech Emotion Recognition (SER) systems is undermined by their fragility to domain shifts, such as speaker variability, the distinction between acted and naturalistic emotions, and cross-corpus variations. While domain adaptation and fine-tuning are widely studied, they require either source data or labelled target data, which are often unavailable or raise privacy concerns in SER. Test-time adaptation (TTA) bridges this gap by adapting models at inference using only unlabeled target data. Yet, having been predominantly designed for image classification and speech recognition, the efficacy of TTA for mitigating the unique domain shifts in SER has not been investigated. In this paper, we present the first systematic evaluation and comparison covering 11 TTA methods across three representative SER tasks. The results indicate that backpropagation-free TTA methods are the most promising. Conversely, entropy minimization and pseudo-labeling generally fail, as their core assumption of a single, confident ground-truth label is incompatible with the inherent ambiguity of emotional expression. Further, no single method universally excels, and its effectiveness is highly dependent on the distributional shifts and tasks.

语音识别情绪识别测试时适应无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。