arXiv:2608.18087cs.CLcs.AI2026-08中稿 · Interspeech 2026

针对印地语等语言的分词缺陷,提出保留词根完整性的新分词方法。

SuTRA : Structurally-Unified Tokenization with Root Awareness

论文配图:SuTRA : Structurally-Unified Tokenization with Root Awareness
图 1 · 摘自论文原文
  • 基于词根意识设计分词算法,避免破坏词语结构
  • 在印地语上提升34%语义还原能力,边界准确率提高14.7%
  • 适用于需要保持形态结构的语言处理任务

现有子词分词器侧重统计压缩,忽略形态结构,尤其忽视词根与词缀的关系。这对以复杂音节(akshara)为基本单位的印地语等形态丰富的语言有害。频率驱动的方法会过度拆分词语,任意切割词根与词缀,这种现象称为形态破碎。我们提出SuTRA(具有词根意识的结构统一分词),一种形态感知算法,能保持akshara完整性,并惩罚跨形态边界合并。同时发布了涵盖印地语、马拉地语和古吉拉特语的新形态分割数据集。SuTRA有效减少破碎,在形态对齐(边界F1)上较BPE提升14.7%,印地语语义可恢复性提升34%。这些结构优势带来机器翻译平均8.08点chrF2提升。

原文摘要 · Abstract (English)

Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering. We propose SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries. We also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati. SuTRA reduces shattering, achieving peak gains of +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability (Hindi) over BPE. These structural gains yield an average improvement of +8.08 chrF2 in machine translation.

分词形态学印地语语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。