arXiv:2605.19069cs.CLcs.AI2026-05被引 1

评测五家商用语音识别系统在多语言切换场景下的表现,发现主流模型对阿拉伯语、波斯语和德语切换识别仍有显著差距。

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

论文配图:Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
图 1 · 摘自论文原文
  • 构建四组多语言切换数据集,用大模型筛选减少91%成本
  • 十一音声学模型在所有语种中最低错误率(13.2%)和最高语义相似度(0.936)
  • 传统词错误率夸大差距,语义评分更真实反映识别质量

代码切换——单个话语中自然交替使用两种语言——仍是自动语音识别(ASR)中最具挑战性且研究不足的场景之一。本文评估了五家商业ASR提供商在四种语言组合上的表现:埃及阿拉伯语-英语、沙特阿拉伯语(纳吉迪/希贾兹)-英语、波斯语(波斯语)-英语以及德语-英语,每组包含300个样本,通过两阶段筛选流程结合启发式过滤与GPT-4o和Gemini 1.5 Pro集成评分器,将大模型成本降低约91%。评估采用词错误率(WER)和BERTScore两种指标,结果显示在所有阿拉伯语和波斯语组合中,两个指标对系统排序一致(τ=1.0),但WER因惩罚语义正确但拼写不同的转录,使质量差距被夸大约3倍。ElevenLabs Scribe v2在总体上取得最低的词错误率(13.2%)并领先于其他系统,在整体上达到0.936的BERTScore。分难度层级分析揭示了聚合平均值掩盖的性能差异,而BERT嵌入投影证实了参考文本与生成结果在语义层面高度接近,尽管表面书写形式不同。数据集已公开发布于https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch。

原文摘要 · Abstract (English)

Code-switching -- the natural alternation between two languages within a single utterance -- remains one of the most challenging and under-studied conditions for automatic speech recognition (ASR). We present a benchmark evaluating five commercial ASR providers across four language pairs: Egyptian Arabic--English, Saudi Arabic (Najdi/Hijazi)--English, Persian (Farsi)--English, and German--English, comprising 300 samples per pair selected by a two-stage pipeline combining heuristic filtering with a GPT-4o and Gemini 1.5 Pro ensemble scorer, reducing LLM costs by $\approx$91\%. We evaluate on both WER and BERTScore, showing that while both metrics agree on the ordinal ranking of systems for all Arabic and Persian pairs ($τ= 1.0$), WER inflates the magnitude of quality gaps by approximately 3$\times$ by penalising semantically correct transliteration choices. ElevenLabs Scribe v2 achieves the lowest WER (13.2\% overall) and leads on BERTScore (0.936 overall). Difficulty-stratified analysis reveals performance gaps masked by aggregate averages, and BERT embedding projections confirm semantic proximity between reference and hypothesis despite surface-level script differences. The dataset is publicly available at https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.

语音识别多语言代码切换评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。