arXiv:2605.27984cs.CLcs.AI2026-05

构建韩语语音模型评估基准,揭示英语评测的局限性。

KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs

论文配图:KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
图 1 · 摘自论文原文
  • 用双框架迁移构建韩语语音问答与音频理解数据集
  • 涵盖12,345个样本,发现英韩性能差距因模型和任务而异
  • 适合多语言语音模型开发者与评测研究者参考

语音语言模型(SpeechLMs)通过将大语言模型扩展到语音模态取得了显著进展。然而,当前语音模型评估仍高度集中于英语,限制了对多语言语音能力的可靠评估。直接通过自动语音识别(ASR)、翻译、归一化和文本转语音(TTS)迁移基准会破坏语言特异性指令、答案约束及口语表达形式;对于音频理解任务,源语言音频的迁移也无法保留目标语言说话人特征、口音和副语言属性。为解决这些问题,我们提出两种基于人类代理的基准构建框架:一种将源语言语音问答(SpokenQA)基准迁移至目标语言,另一种利用转录文本和说话人元数据,将目标语言的ASR语料转化为音频理解基准。基于此,我们构建并公开发布三个韩语语音基准:用于韩语语音问答的KVoiceBench和KOpenAudioBench,以及用于韩语音频理解的KMMAU,共计包含12,345个样本。我们评估了八种近期的SpeechLMs,发现英韩性能差距在不同模型和任务类别间差异显著,且语音问答与音频理解的排名存在分化,暴露出英语单一评估无法发现的互补性缺陷。

原文摘要 · Abstract (English)

Speech language models (SpeechLMs) have achieved substantial progress by extending large language models (LLMs) to the speech modality. However, SpeechLM evaluation remains heavily centered on English, limiting reliable assessment of multilingual speech capabilities. Straightforward benchmark transfer through ASR, translation, normalization, and TTS can corrupt language-specific instructions, answer constraints, and spoken forms; for audio understanding, transferring source-language audio also fails to preserve target-language speaker attributes, accents, and paralinguistic properties. To address these limitations, we propose two human-agent benchmark-construction frameworks: one transfers source-language SpokenQA benchmarks into target-language SpokenQA benchmarks, and the other converts target-language ASR corpora into audio understanding benchmarks using transcriptions and speaker metadata. Using these frameworks, we construct and publicly release three Korean speech benchmarks: KVoiceBench and KOpenAudioBench for Korean SpokenQA, and KMMAU for Korean audio understanding, comprising 12,345 samples in total. We evaluate eight recent SpeechLMs and find that English-Korean performance gaps vary substantially across models and task families, and that SpokenQA and audio understanding rankings diverge, revealing complementary weaknesses invisible to English-only evaluation.

语音评测多语言韩国语基准构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。