arXiv:2602.03873cs.SDcs.AI2026-02被引 3

用测试时缩放提升语音情感模型对模糊情绪的识别能力

Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models

  • 引入测试时缩放技术增强音频语言模型对模糊情绪的推理能力
  • 在三个数据集上验证,显著提升模型在模糊情绪场景下的准确率
  • 为构建更懂人类情感的对话系统提供新方向,适合情感计算研究者

从语音中识别情绪是实现社交感知对话AI的关键。然而,以往研究多将情绪识别视为分类问题,而现实中的情感状态常具模糊性、重叠性和情境依赖性,给标注和建模带来挑战。近期的大规模音频语言模型(ALMs)虽能在无显式情感标注下进行细致的情感推理,但其处理模糊情绪的能力仍待探索。同时,测试时缩放(TTS)等推理阶段技术在复杂NLP任务中展现出提升泛化与适应性的潜力,但在情感计算中的应用尚不明确。本文首次构建了在测试时缩放下针对语音模糊情绪识别的基准。系统评估了八种先进ALMs与五种TTS策略在三个主流语音情感数据集上的表现。深入分析了模型能力、TTS与情感模糊性之间的交互关系,揭示了模糊情感理解中的计算与表征挑战。该基准为开发更鲁棒、情境敏感且具备情感智能的语音系统奠定基础,并指明未来关键研究方向。

原文摘要 · Abstract (English)

Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classification problem, real-world affective states are often ambiguous, overlapping, and context-dependent, posing significant challenges for both annotation and automatic modeling. Recent large-scale audio language models (ALMs) offer new opportunities for nuanced affective reasoning without explicit emotion supervision, but their capacity to handle ambiguous emotions remains underexplored. At the same time, advances in inference-time techniques such as test-time scaling (TTS) have shown promise for improving generalization and adaptability in hard NLP tasks, but their relevance to affective computing is still largely unknown. In this work, we introduce the first benchmark for ambiguous emotion recognition in speech with ALMs under test-time scaling. Our evaluation systematically compares eight state-of-the-art ALMs and five TTS strategies across three prominent speech emotion datasets. We further provide an in-depth analysis of the interaction between model capacity, TTS, and affective ambiguity, offering new insights into the computational and representational challenges of ambiguous emotion understanding. Our benchmark establishes a foundation for developing more robust, context-aware, and emotionally intelligent speech-based AI systems, and highlights key future directions for bridging the gap between model assumptions and the complexity of real-world human emotion.

情感识别音频语言模型测试时缩放模糊情绪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。