通过分析令牌级逻辑值,揭示语音大模型在模糊情绪识别中的潜在能力。
Token-Level Logits Matter: A Closer Look at Speech Foundation Models for Ambiguous Emotion Recognition
- 利用令牌级逻辑值解析模型内部情绪推理过程
- 模型虽文本输出不一致,但内在情绪判断具鲁棒性
- 适合研究情感认知机制与模型可解释性的学者
对话式人工智能中的情感智能在人机交互等领域至关重要。尽管已有众多模型被开发,但往往忽视人类情感固有的复杂性与模糊性。在大规模语音基础模型(SFM)时代,理解其对模糊情绪的识别能力对下一代情感感知模型的发展尤为关键。本研究考察了SFM在模糊情绪识别中的表现,设计了用于模糊情绪预测的提示,并提出两种新方法推断模糊情绪分布:一种基于生成文本响应,另一种通过分析模型内部的令牌级逻辑值。结果表明,尽管SFM在模糊情绪下的文本生成并不稳定,但其可通过先验知识在令牌层面准确理解情绪,表现出对不同提示的强鲁棒性。
原文摘要 · Abstract (English)
Emotional intelligence in conversational AI is crucial across domains like human-computer interaction. While numerous models have been developed, they often overlook the complexity and ambiguity inherent in human emotions. In the era of large speech foundation models (SFMs), understanding their capability in recognizing ambiguous emotions is essential for the development of next-generation emotion-aware models. This study examines the effectiveness of SFMs in ambiguous emotion recognition. We designed prompts for ambiguous emotion prediction and introduced two novel approaches to infer ambiguous emotion distributions: one analysing generated text responses and the other examining the internal processing of SFMs through token-level logits. Our findings suggest that while SFMs may not consistently generate accurate text responses for ambiguous emotions, they can interpret such emotions at the token level based on prior knowledge, demonstrating robustness across different prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。