arXiv:2604.15929cs.CL2026-04

构建多语言科研对话数据集,评测ASR系统跨语言混合输入能力

MUSCAT: MUltilingual, SCientific ConversATion Benchmark

论文配图:MUSCAT: MUltilingual, SCientific ConversATion Benchmark
图 1 · 摘自论文原文
  • 设计双语科研对话场景,模拟真实跨语言交流
  • 现有顶尖ASR在多语言混合输入下表现仍不理想
  • 提供超越WER的评估框架,适合多语言语音研究者

多语言语音技术的目标是实现不同语言使用者间的无缝沟通,使人人都像多语者一样自然交流。为达成此目标,语音技术需应对混合多语言输入、专业术语及语码转换等挑战。然而当前缺乏针对此类情境的数据集与基准测试。本文提出MUSCAT基准,用于评估自动语音识别(ASR)系统处理这些挑战的能力。该基准包含多位发言者以不同语言进行的双语科研论文讨论,涵盖真实语码转换场景。我们提供超越词错误率(WER)的标准评估框架,支持跨语言性能一致比较。实验表明,当前最先进的ASR系统在该数据集上仍面临显著挑战。数据集已公开于https://huggingface.co/datasets/goodpiku/muscat-eval。

原文摘要 · Abstract (English)

The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experience as though everyone were a multilingual speaker. To create this experience, speech technology needs to address several challenges: Handling mixed multilingual input, specific vocabulary, and code-switching. However, there is currently no dataset benchmarking this situation. We propose a new benchmark to evaluate current Automatic Speech Recognition (ASR) systems, whether they are able to handle these challenges. The benchmark consists of bilingual discussions on scientific papers between multiple speakers, each conversing in a different language. We provide a standard evaluation framework, beyond Word Error Rate (WER) enabling consistent comparison of ASR performance across languages. Experimental results demonstrate that the proposed dataset is still an open challenge for state-of-the-art ASR systems. The dataset is available in https://huggingface.co/datasets/goodpiku/muscat-eval. Keywords: multilingual, speech recognition, audio segmentation, speaker diarization

多语言语音识别对话数据集跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。