arXiv:2606.20993cs.CL2026-06ACL被引 1

用国际音标统一多语言分词,提升非拉丁语系表现

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet

论文配图:Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet
图 1 · 摘自论文原文
  • 以国际音标(IPA)作为跨语言输入表示,减少分词偏差
  • 24种语言、14种文字测试中,非拉丁语系分词质量显著提升
  • 适合关注多语言模型公平性与泛化能力的研究者

多语言模型在不同语言间常出现性能差异,问题可能早在分词阶段就已出现。主流子词分词方法偏重高资源语言,而无分词器方法对字符字节比高的书写系统仍产生更长序列。为此,我们提出使用国际音标(IPA)作为多语言分词器的通用输入表示。IPA具有紧凑符号集、更高的跨语言字符重叠度,以及更均衡的字节/字符分布。我们在24种语言、14种书写系统上训练了文本与IPA子词分词器的对应对,并证明IPA分词器在非拉丁语系中持续提升分词质量,且对未见语言和书写系统具有更强泛化能力。

原文摘要 · Abstract (English)

Multilingual language models often exhibit performance disparities across languages that can arise as early as the tokenization stage. Widely-used subword tokenization approaches favor high-resource languages, and tokenizer-free methods still yield longer sequences for scripts with a higher bytes-per-character ratio. To address these shortcomings, we propose to use the International Phonetic Alphabet (IPA) as a language-agnostic input representation for multilingual tokenizers. IPA provides a compact symbol inventory, greater cross-lingual character overlap, and a more balanced byte-per-character distribution across languages. We train matched pairs of text vs. IPA subword tokenizers across 24 languages and 14 scripts and demonstrate that IPA tokenizers consistently improve tokenization quality, especially for non-Latin scripts, and generalize more effectively to unseen languages and scripts.

多语言分词IPA跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。