融合语音与文本特征,提升芬兰语情感判断准确率
Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability

- 联合语音与文本特征建模情感感知
- 文本+语音组合提升愉悦度预测效果
- 适用于自然口语情感分析研究者
现有语音情感研究显示,语音的声学特征和语义内容均与感知到的情感唤醒度和愉悦度相关。但两者在感知过程中的相对贡献尚不明确,尤其对芬兰语而言,多数研究仅聚焦声学或文本分析。本文利用新发布的自然芬兰语情感语音语料库,系统探讨文本与音频特征在建模人类情感愉悦度与唤醒度感知中的协同作用。结果表明,文本与音频特征结合可显著提升愉悦度回归性能,而唤醒度回归中互补效应不明显。研究结果支持其他语言的发现,为自然芬兰语情感分析提供了新数据与认知依据。
原文摘要 · Abstract (English)
Existing research on affect in speech has shown how acoustic surface characteristics and content-related linguistic aspects of speech both relate to perceived emotional arousal and valence. However, it is not clear what the relative contributions of these two factors are in the perceptual process. This is especially true for Finnish, for which most existing studies focus on either acoustic-phonetic or text analysis. This paper presents a study where we systematically explore the combinatory role of text- and audio-based features in modeling the human perception of valence and arousal using a newly released affective speech corpus for spontaneous Finnish. We show that the combination of text- and audio-based features improves valence regression results over the individual modalities, whereas for arousal regression the complementary effect is not substantial. The results support prior findings from other languages, providing new data and knowledge on spontaneous Finnish speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。