arXiv:2607.17812cs.CL2026-07

首个面向复杂环境的西班牙语语音理解基准,评测大模型真实场景表现

ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions

论文配图:ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions
图 1 · 摘自论文原文
  • 构建162.9小时真实录音,覆盖多口音与非标准发音
  • 包含10类推理题,测试模型在噪声等复杂条件下的理解能力
  • 适合评估大语音模型在真实世界应用中的鲁棒性

随着大规模音频语言模型(LALMs)的发展,鲁棒的评估框架变得至关重要。然而,针对真实声学条件下西班牙语语音理解的研究仍十分匮乏。本文提出ESCUCHA,首个专为评估LALMs在异构声学条件和推理能力下表现而设计的西班牙语语音理解基准。ESCUCHA包含1,000个由人工标注的问题与对应音频,总时长达162.9小时,数据均来自真实场景而非现有数据集,单段时长从几秒到超过80分钟不等。该基准强调推理能力,涵盖9类感知任务与10类推理任务,并通过多种西班牙语口音及非标准发音体现语言多样性。此外,还包含多音频问题、口语提问和音频指令,并标记支持开放式评估的问题。对多个先进多模态与语音模型的基准测试显示,其性能与受训人类之间存在显著差距。

原文摘要 · Abstract (English)

As large audio language models (LALMs) advance, robust evaluation frameworks have become essential. In this context, Spanish speech understanding under realistic acoustic conditions has received particularly little attention. We introduce ESCUCHA, the first Spanish speech understanding benchmark designed to evaluate LALMs across heterogeneous acoustic conditions and reasoning abilities. ESCUCHA comprises 1,000 human-curated questions paired with audio, totaling 162.9 hours sourced directly ``from the wild'' rather than drawn from existing datasets, with durations ranging from a few seconds to over 80 minutes. The benchmark emphasizes reasoning, spanning 9 perceptual and 10 reasoning categories, and it captures linguistic diversity through multiple Spanish accents and non-normative speech. ESCUCHA further includes multi-audio questions, spoken questions, and audio instructions, and it flags which questions support open-ended evaluation. Benchmarking several state-of-the-art multimodal and speech models reveals substantial performance gaps relative to trained humans.

语音理解多语种基准测试大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。