33M小模型实现高精度语音转音标,媲美大模型表现
BranchShine: Compact Raw-Audio-to-IPA Transcription with a RoPE E-Branchformer Encoder

- 用轻量卷积前端+19层RoPE E-Branchformer编码器,直接处理原始音频
- 在41种语言上达到9.19%的音标字符错误率,优于5.75亿参数基线
- 适合资源受限场景,尤其擅长避免误读,适合儿童语音分析
语音转音标对发音研究很有价值,但现有多语言系统通常庞大且评估受归一化方式影响。本文提出BranchShine,一个3300万参数的原始音频到音标CTC识别器,采用轻量卷积前端和19层RoPE E-Branchformer编码器。在包含41个语言标签的16,660条语句多语言测试集上,分支结构在匹配归一化条件下取得9.19%的空白无关音标字符错误率,优于5.75亿参数的PhoneticXEUS基线(9.78%)。对儿童阅读的二次分析显示:BranchShine在错误读音上更保守,而Whisper-Medium在正确读音接受率上更强。结果表明,紧凑的原始音频到音标模型可在字符级音标转写上接近大型模型性能。
原文摘要 · Abstract (English)
Speech-to-IPA transcription is useful when the desired output is pronunciation rather than orthographic text, but competitive multilingual systems are often large and evaluation is sensitive to normalization choices. This paper presents BranchShine, a 33M-parameter raw-audio CTC recognizer with a lightweight convolutional front end and a 19-block RoPE E-Branchformer encoder. We find that BranchShine provides a compact and competitive operating point for IPA transcription under matched normalization and scoring. On a 16,660-utterance multilingual test set covering 41 language labels, BranchShine obtains 9.19% whitespace-insensitive IPA character error rate, compared with 9.78% for the 575.00M-parameter PhoneticXEUS baseline. A secondary child speech reading analysis shows a complementary operating profile: BranchShine is more conservative on incorrect readings, while Whisper-Medium is stronger on exact acceptance of correct readings. Overall, the results indicate that a compact raw-audio-to-IPA model can approach much larger baselines on character-level IPA transcription.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。