arXiv:2412.09102cs.CLcs.AI2024-12被引 4

多语言音素转字形模型,提升名字音译准确率。

PolyIPA -- Multilingual Phoneme-to-Grapheme Conversion Model

  • 用音素转字形技术实现跨语言名字音译
  • 字符错误率仅0.055,顶3候选使误差降低52.7%
  • 适合多语言音译、姓名研究与信息检索场景

本文提出PolyIPA,一种用于多语言名字音译、专名学研究和信息检索的新型多语言音素转字形模型。该模型利用两个辅助模型进行数据增强:IPA2vec用于跨语言寻找发音相似词,similarIPA用于处理音标表示差异。在涵盖多种语言和书写系统的测试集上,模型取得平均字符错误率0.055和字符级BLEU分数0.914,尤其在浅层正字法语言上表现优异。采用束搜索后,前3个候选结果将有效错误率降低52.7%(降至CER: 0.026),充分展现其在跨语言应用中的有效性。

原文摘要 · Abstract (English)

This paper presents PolyIPA, a novel multilingual phoneme-to-grapheme conversion model designed for multilingual name transliteration, onomastic research, and information retrieval. The model leverages two helper models developed for data augmentation: IPA2vec for finding soundalikes across languages, and similarIPA for handling phonetic notation variations. Evaluated on a test set that spans multiple languages and writing systems, the model achieves a mean Character Error Rate of 0.055 and a character-level BLEU score of 0.914, with particularly strong performance on languages with shallow orthographies. The implementation of beam search further improves practical utility, with top-3 candidates reducing the effective error rate by 52.7\% (to CER: 0.026), demonstrating the model's effectiveness for cross-linguistic applications.

音素转字形多语言名字音译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。