arXiv:2509.25495cs.SDcs.AI2025-09中稿 · ICASSP 2026被引 4

无需训练,通过统计更新提升语音情感识别在分布外场景的性能

EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition

  • 利用期望最大化算法动态更新类别条件统计量,实现测试时分布估计
  • 在6个分布外基准上均超越现有方法,平均提升超过3个百分点
  • 轻量无参设计,适合实时部署,尤其适用于资源受限场景

基于音频-语言模型(ALMs)的语音情感识别(SER)在测试时面对分布偏移仍易退化。测试时自适应(TTA)虽有潜力,但常依赖梯度更新或提示调优,灵活性不足。本文提出Emo-TTA,一种轻量、无需训练的自适应框架,通过期望最大化过程对类条件统计量进行增量更新,以模型预测为先验,实现显式的测试时分布估计。该方法仅作用于单个测试样本,不修改模型权重。在六个分布外SER基准上的实验表明,相比现有TTA基线,本方法实现了持续的准确率提升,验证了统计自适应在对齐模型输出与动态测试分布方面的有效性。

原文摘要 · Abstract (English)

Speech emotion recognition (SER) with audio-language models (ALMs) remains vulnerable to distribution shifts at test time, leading to performance degradation in out-of-domain scenarios. Test-time adaptation (TTA) provides a promising solution but often relies on gradient-based updates or prompt tuning, limiting flexibility and practicality. We propose Emo-TTA, a lightweight, training-free adaptation framework that incrementally updates class-conditional statistics via an Expectation-Maximization procedure for explicit test-time distribution estimation, using ALM predictions as priors. Emo-TTA operates on individual test samples without modifying model weights. Experiments on six out-of-domain SER benchmarks show consistent accuracy improvements over prior TTA baselines, demonstrating the effectiveness of statistical adaptation in aligning model predictions with evolving test distributions.

语音情感识别测试时自适应无参优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。