arXiv:2604.08786cs.SDeess.AS2026-04被引 1

提出新指标SFR,发现多语种语音识别常错用文字系统。

Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmark

论文配图:Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmark
图 1 · 摘自论文原文
  • 定义无需参考文本的脚本保真率SFR,衡量输出字符是否在目标书写系统内
  • 100组模型-语言对中21组出现脚本崩溃(SFR<10%),主要集中在Whisper模型
  • 通过提示词设计可显著提升脚本保真度,尤其对乌尔都语等语言效果明显

词错误率(WER)是语音识别主流评估指标,但无法检测一种系统性失败:模型生成流畅却使用错误书写系统的输出。本文定义了无需参考转录的脚本保真率(SFR),即假设文本中属于目标书写系统区块的字符比例。在涵盖六种书写系统、十种语言和十种模型(七种Whisper尺寸、MMS-1B、SeamlessM4T-v2、Gemma 4 E2B)的FLEURS测试集上进行系统测量。100个模型-语言组合中,21个(21%;95%置信区间:14-30%)出现脚本崩溃(SFR低于10%),其中20个涉及Whisper,1个为Gemma 4 E2B在乌尔都语上使用通用提示词时发生。在十语言的Gemma 4探测实验中,引入脚本感知提示后,平均SFR从71.2%提升至97.7%,修复乌尔都语崩溃(6.5%→97.0%),并使六个基础SFR低于90%的语言在下游NLLB翻译任务中恢复5.9点chrF得分。识别出四种崩溃模式:拉丁音似替换、索马里语中阿拉伯字母替代、印地语/马拉雅拉姆语中天城文替代,以及格鲁吉亚语特有的独特脚本拉丁化现象。

原文摘要 · Abstract (English)

Word error rate (WER) is the dominant metric for automatic speech recognition, yet it cannot detect a systematic failure mode: models that produce fluent output in the wrong writing system. We define Script Fidelity Rate (SFR), the fraction of hypothesis characters in the target script block, computable without reference transcriptions, and report a systematic measurement of script collapse across ten languages spanning six writing systems and ten models (seven Whisper sizes, MMS-1B, SeamlessM4T-v2, and Gemma 4 E2B) on FLEURS test sets. Across 100 evaluated model-language pairs, 21 (21%; 95% Wilson CI: 14-30%) exhibit script collapse (SFR less than 10%): 20 involve Whisper and one involves Gemma 4 E2B on Urdu under a generic transcription prompt. In a ten-language Gemma 4 probe, script-aware prompting raises mean SFR from 71.2% to 97.7%, fixes Urdu collapse (6.5% to 97.0%), and recovers 5.9 chrF on downstream NLLB translation for the six languages whose baseline SFR is below 90%. We identify four collapse patterns: Latin phonetic substitution, Arabic substitution for Somali, Devanagari substitution for Bengali/Malayalam, and unique-script Latin collapse for Georgian.

语音识别多语言脚本保真评测指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。