构建110种语言方言的语音理解基准,评估模型真实语义理解能力。
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

- 融合真人录音与指令生成合成语音,覆盖19种中文方言和80多种低资源语言。
- 开源模型在方言上表现优于传统分步系统,但对低资源语言性能严重下降。
- 发现零样本提示常降低性能,暴露当前模型在语音-文本对齐上的缺陷。
尽管端到端语音大模型快速发展,其评估仍停留在简单转录阶段。现有基准存在三大局限:严重偏向高资源语言、聚焦低层识别(如语音识别)而非语义推理、忽视地区方言。为此,我们推出PolySpeech-100,一个大规模基准,用于评估110种语言变体的「母语级」语音理解能力。采用新型混合构建流程,结合高质量人工录音与指令驱动的合成语音,覆盖19种中文方言及超过80种低资源语言。对22个前沿模型(包括Gemini-3、GPT-Audio、Qwen2.5-Omni)的评估揭示:第一,开源端到端模型在复杂方言上优于分步系统(ASR+LLM),表明直接音频处理更保留音调、重音等关键韵律特征;第二,商业模型保持鲁棒性,而开源模型在低资源语言上出现灾难性退化;第三,在标准零样本设置下,多数模型的思维链提示反而降低性能,暴露当前架构中模态对齐的潜在缺陷。数据、演示与代码已公开于https://github.com/YoungSeng/PolySpeech-100。
原文摘要 · Abstract (English)
While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription. Existing benchmarks suffer from three critical limitations: a pronounced bias towards high-resource languages, a focus on low-level recognition (ASR) rather than semantic reasoning, and a neglect of regional dialects. To bridge this gap, we introduce PolySpeech-100, a massive-scale benchmark designed to assess `native-level' speech comprehension across 110 linguistic variants. We employ a novel hybrid construction pipeline that augments gold-standard human recordings with instruction-driven synthetic speech, allowing us to cover 19 distinct Chinese dialects and over 80 low-resource languages. Extensive evaluation of 22 state-of-the-art models (including Gemini-3, GPT-Audio, and Qwen2.5-Omni) yields pivotal insights. First, we demonstrate that open-source E2E models outperform Cascade (ASR+LLM) systems on heavy dialects, proving that direct audio processing preserves critical paralinguistic cues and prosodic features (e.g., intonation, stress) that are often lost in standard transcription. Second, we reveal a significant performance gap: while commercial models maintain robustness, open-source models suffer catastrophic degradation on low-resource languages. Finally, counter-intuitively, we observe that under standard zero-shot settings, Chain-of-Thought prompting frequently degrades speech understanding performance for most evaluated models, revealing a potential modality alignment gap in current architectures. PolySpeech-100 establishes a rigorous standard for the next generation of inclusive, omni-capable Speech-LLMs. The data, demo, and code are publicly available at https://github.com/YoungSeng/PolySpeech-100.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。