构建语音理解基准HPSU,评估模型是否具备人类级听觉感知能力。
HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
- 融合多模态信息实现高效精准的半自动标注
- 覆盖2万+样本,涵盖从说话人属性到隐含意图的全谱任务
- 揭示当前顶尖模型在真实对话理解上仍远逊人类
近年来,语音大语言模型(Speech LLMs)在自动语音识别(ASR)和语音情感识别(SER)等任务中取得显著进展。然而,这些模型能否达到人类级听觉感知水平,尤其是在理解真实口语中隐含意图与微妙情绪方面,仍缺乏系统评估。为此,我们提出人类级语音理解感知基准HPSU,涵盖超过20,000个经专家验证的中英文语音理解样本。该基准建立了一个涵盖从基础说话人属性识别到复杂隐含意图与情绪推理的综合性评估框架。为应对真实场景下数据稀缺与人工标注成本高的问题,我们开发了融合音频、文本与视觉信息的半自动标注流程,显著提升标注效率与质量。我们系统评估了多个开源及专有语音大模型,结果表明,即使是最先进的模型,在理解真实口语交互方面仍明显落后于人类。因此,HPSU将为推动语音大模型向人类级感知与认知发展提供关键支持。
原文摘要 · Abstract (English)
Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of their ability to comprehend latent intentions and implicit emotions in real-world spoken language, remains underexplored. To this end, we introduce the Human-level Perception in Spoken Speech Understanding (HPSU), a new benchmark for fully evaluating the human-level perceptual and understanding capabilities of Speech LLMs. HPSU comprises over 20,000 expert-validated spoken language understanding samples in English and Chinese. It establishes a comprehensive evaluation framework by encompassing a spectrum of tasks, ranging from basic speaker attribute recognition to complex inference of latent intentions and implicit emotions. To address the issues of data scarcity and high cost of manual annotation in real-world scenarios, we developed a semi-automatic annotation process. This process fuses audio, textual, and visual information to enable precise speech understanding and labeling, thus enhancing both annotation efficiency and quality. We systematically evaluate various open-source and proprietary Speech LLMs. The results demonstrate that even top-performing models still fall considerably short of human capabilities in understanding genuine spoken interactions. Consequently, HPSU will be useful for guiding the development of Speech LLMs toward human-level perception and cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。