arXiv:2604.19782cs.CLcs.AI2026-04

首个针对韩语语音理解与忠实度的综合评测基准

KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness

论文配图:KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
图 1 · 摘自论文原文
  • 构建六项任务,覆盖语音理解与忠实度评估
  • 引入韩国高考听力题和本土文化内容提升评测真实性
  • 支持白盒与黑盒模型评测,公开代码与榜单

大型音频语言模型(LALMs)在多语言语音理解方面取得进展,但非英语语言的评测基准仍严重不足,韩语即为典型代表。本文提出KoALa-Bench,一个面向韩语语音理解与语音忠实度的综合性评测基准。该基准包含六项任务:四项评估基础语音理解能力(自动语音识别、语音翻译、语音问答、语音指令遵循),两项评估语音忠实度——源于观察发现多数LALMs未能充分利用语音模态。为体现韩国本土知识,基准融入韩国大学学力水平考试听力题及韩国文化相关内容。我们在六种模型(含白盒与黑盒)上开展全面实验,评测数据集、代码及排行榜已公开于https://ksbench.github.io/Korean-Benchmark/。

原文摘要 · Abstract (English)

Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexplored case. In this paper, we introduce KoALa-Bench, a comprehensive benchmark for evaluating Korean speech understanding and speech faithfulness of LALMs. In particular, KoALa-Bench comprises six tasks. Four tasks evaluate fundamental speech understanding capabilities, including automatic speech recognition, speech translation, speech question answering, and speech instruction following, while the remaining two tasks evaluate speech faithfulness, motivated by our observation that several LALMs often fail to fully leverage the speech modality. Furthermore, to reflect Korea-specific knowledge, our benchmark incorporates listening questions from the Korean college scholastic ability test as well as content covering Korean cultural domains. We conduct extensive experiments across six models, including both white-box and black-box ones. Our benchmark, evaluation code, and leaderboard are publicly available at https://ksbench.github.io/Korean-Benchmark/.

语音理解韩语评测基准LALM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。