构建语音推理评测基准,揭示大模型听懂话但未必能推理的短板
SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
- 设计三维度评测框架:事实检索、流程推断、规范判断
- 11个顶尖模型测试显示高转录准确率不等于强推理能力
- 支持多选、生成、声学特征三种评估形式,覆盖真实对话场景
大型音频-语言模型(LALMs)在句子级转录和情感识别上已接近人类水平。然而,现有评估主要聚焦表面感知任务,对模型在语音场景中进行上下文理解和推理的能力缺乏充分考察。为填补这一空白,我们提出SpeechR,一个统一的语音推理评测基准。该基准从三个关键维度评估模型:事实检索、流程推断与规范判断,并包含三种不同评估形式:多选题衡量答案选择准确性,生成式任务评估推理链的连贯性与逻辑一致性,声学特征测试则探究语调与情绪变化对推理表现的影响。对11个前沿LALM的测评表明,高转录准确率并不等同于强推理能力。SpeechR为语音语言模型的推理能力评估提供了结构化框架,有助于更精准地分析模型在多样化对话任务中的表现。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for contextual and inference-driven reasoning in speech-based scenarios insufficiently examined. To address this gap, we introduce SpeechR, a unified benchmark for evaluating reasoning over speech in large audio-language models. SpeechR evaluates models along three key dimensions: factual retrieval, procedural inference, and normative judgment. It includes three distinct evaluation formats. The multiple-choice version measures answer selection accuracy. The generative version assesses the coherence and logical consistency of reasoning chains. The acoustic-feature version investigates whether variations in stress and emotion affect reasoning performance. Evaluations on eleven state-of-the-art LALMs reveal that high transcription accuracy does not translate into strong reasoning capabilities. SpeechR establishes a structured benchmark for evaluating reasoning in spoken language, enabling more targeted analysis of model capabilities across diverse dialogue-based tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。