arXiv:2507.19361cs.CLcs.AI2025-07ACL被引 3

用认知分级评估大模型语音理解能力,发现传统指标的局限

SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models

  • 按布鲁姆认知分类设计三层评估:记忆、理解、应用
  • 揭示现有基准标注错误和大模型语音幻觉问题
  • 适合研究语音-语言融合模型与评测方法的学者

我们提出基于语音的智商(SIQ)评估框架,用于衡量大语言模型在语音理解任务中的认知级表现。该框架借鉴布鲁姆认知分类学,从三个层次评估:(1) 记忆层(以词错误率WER衡量原文准确性);(2) 理解层(比较模型对语音内容的语义相似性);(3) 应用层(模拟下游任务的问答准确率)。实验表明,SIQ不仅能量化模型语音理解能力,还能统一比较级联式方法(如ASR+LLM)与端到端模型,识别现有基准中的标注错误,并检测大模型在语音理解中的幻觉现象。该框架首次将认知理论与语音评测结合,揭示了多模态训练中被忽视的挑战。代码与数据将开源,促进后续研究。

原文摘要 · Abstract (English)

We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, designed to assess their voice understanding ability. Moving beyond popular voice understanding metrics such as word error rate (WER), SIQ examines LLM Voice across three cognitive levels motivated by Bloom's Taxonomy: (1) Remembering (i.e., WER for verbatim accuracy); (2) Understanding (i.e., similarity of LLM's interpretations); and (3) Application (i.e., QA accuracy for simulating downstream tasks). We demonstrate that SIQ not only quantifies voice understanding abilities but also provides unified comparisons between cascaded methods (e.g., ASR LLM) and end-to-end models, identifies annotation errors in existing benchmarks, and detects hallucinations in LLM Voice. Our framework represents a first-of-its-kind intelligence examination that bridges cognitive principles with voice-oriented benchmarks, while exposing overlooked challenges in multi-modal training. Our code and data will be open source to encourage future studies.

语音理解认知评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。