arXiv:2509.16329eess.AScs.SD2025-09中稿 · APSIPA-ASC 2025

多语言语音模型更擅长从嘈杂人群声中识别集体情绪。

Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds

  • 用多语言预训练的语音基础模型处理人群情绪识别
  • 在1秒到250毫秒短音频上均优于单语模型
  • 适合需要快速响应的实时情绪分析场景

本文研究多语言(Polyglot)语音基础模型(SFM)在人群情绪识别(CER)中的应用。我们假设,经过多种语言、口音和语音模式预训练的多语言SFM,能更好应对人群场景中嘈杂复杂的声学环境,从而在CER任务中具备显著优势。为验证该假设,我们在基准CER数据集上进行系统性对比实验,涵盖多语言、单语言及说话人识别型SFM,并测试不同音频长度(1秒、500毫秒、250毫秒)下的表现。结果一致显示,多语言SFM在所有时长下均优于其他模型,甚至在极短输入下仍表现优异。这些发现为推动语音基础模型在CER领域建立新基准提供了重要支持。

原文摘要 · Abstract (English)

This paper investigates the polyglot (multilingual) speech foundation models (SFMs) for Crowd Emotion Recognition (CER). We hypothesize that polyglot SFMs, pre-trained on diverse languages, accents, and speech patterns, are particularly adept at navigating the noisy and complex acoustic environments characteristic of crowd settings, thereby offering a significant advantage for CER. To substantiate this, we perform a comprehensive analysis, comparing polyglot, monolingual, and speaker recognition SFMs through extensive experiments on a benchmark CER dataset across varying audio durations (1 sec, 500 ms, and 250 ms). The results consistently demonstrate the superiority of polyglot SFMs, outperforming their counterparts across all audio lengths and excelling even with extremely short-duration inputs. These findings pave the way for adaptation of SFMs in setting up new benchmarks for CER.

语音模型情绪识别多语言人群分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。