arXiv:2504.16283cs.LG2025-04

现有语音情绪模型对非典型语音泛化能力差,易误判情绪。

Affect Models Have Weak Generalizability to Atypical Speech

  • 测试三种非典型语音特征:发音不清、语调单一、嗓音粗糙
  • 非典型语音下悲伤预测比例显著升高,最高超典型语音三倍
  • 伪标签微调可提升对非典型语音的识别,且不影响正常语音表现

语音和声学特征的变化会影响非典型语音人群的语音情感识别性能。我们在一个非典型语音数据集上评估了公开可用的情感识别模型,对比了典型语音数据集的结果。研究涵盖三类语音异常:发音清晰度(intelligibility)、语调单调性(monopitch)和嗓音粗糙度(harshness)。分析包括:(1) 非典型语音中分类情绪分布趋势;(2) 与典型语音数据集的情绪预测分布对比;(3) 自发语音中文本与语音预测在愉悦度和唤醒度上的相关性。结果表明,非典型语音显著影响情绪模型输出。例如,所有类型和程度的非典型语音中,被预测为悲伤的比例均显著高于典型语音数据集。初步实验发现,在伪标签非典型语音数据上微调模型,可在不损害典型语音性能的前提下提升其在非典型语音上的表现。研究强调需扩大训练与评估数据集覆盖范围,并发展对语音差异具有鲁棒性的建模方法。

原文摘要 · Abstract (English)

Speech and voice conditions can alter the acoustic properties of speech, which could impact the performance of paralinguistic models for affect for people with atypical speech. We evaluate publicly available models for recognizing categorical and dimensional affect from speech on a dataset of atypical speech, comparing results to datasets of typical speech. We investigate three dimensions of speech atypicality: intelligibility, which is related to pronounciation; monopitch, which is related to prosody, and harshness, which is related to voice quality. We look at (1) distributional trends of categorical affect predictions within the dataset, (2) distributional comparisons of categorical affect predictions to similar datasets of typical speech, and (3) correlation strengths between text and speech predictions for spontaneous speech for valence and arousal. We find that the output of affect models is significantly impacted by the presence and degree of speech atypicalities. For instance, the percentage of speech predicted as sad is significantly higher for all types and grades of atypical speech when compared to similar typical speech datasets. In a preliminary investigation on improving robustness for atypical speech, we find that fine-tuning models on pseudo-labeled atypical speech data improves performance on atypical speech without impacting performance on typical speech. Our results emphasize the need for broader training and evaluation datasets for speech emotion models, and for modeling approaches that are robust to voice and speech differences.

语音情感非典型语音模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。