arXiv:2509.16589cs.CLcs.AI2025-09EMNLP被引 9

评测语音大模型理解情绪与语境能力,发现现有模型存在明显短板。

Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data

  • 构建新基准CP-Bench,评估语音大模型结合语言与非语言线索的推理能力
  • 顶尖模型在情感与语境理解任务上准确率不足60%,暴露关键缺陷
  • 适合关注语音智能、人机共情与多模态理解的研究者

近期语音大模型在转录和翻译等任务中表现优异,但在理解语音中的副语言特征(如情感与语调)方面仍显不足,而这对于社交与情感智能至关重要。我们提出CP-Bench,一个用于评估语音大模型在上下文感知的副语言推理能力上的基准,该能力需整合语言内容与非语言线索。基准包含两个精心设计的问答数据集,要求模型具备语言理解与共情能力。我们评估了来自开源与闭源的多个先进语音大模型,并对不同问题类型进行了全面分析。进一步对表现最优的两模型进行温度调节实验,以探究其对任务表现的影响。结果揭示了现有评估体系的关键缺失,并为构建更具上下文感知与情感智能的语音大模型提供了洞见。

原文摘要 · Abstract (English)

Recent speech-LLMs have shown impressive performance in tasks like transcription and translation, yet they remain limited in understanding the paralinguistic aspects of speech crucial for social and emotional intelligence. We propose CP-Bench, a benchmark for evaluating speech-LLMs on contextual paralinguistic reasoning the integration of verbal content with non-verbal cues like emotion and prosody. The benchmark includes two curated question answering (QA) datasets requiring both linguistic and empathetic understanding. We evaluate state-of-the-art speech-LLMs from both open and closed-source models and perform a comprehensive analysis across different question types. The top two models were further analyzed under temperature tuning to understand its effect on this task. Our benchmark reveals a key gap in existing evaluations and offers insights into building more context-aware and emotionally intelligent speech-capable LLMs.

语音大模型情感理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。