arXiv:2511.09197cs.CL2025-11被引 2

让模型动态学习分词,提升复杂语言的生成与跨语言迁移能力

The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages

  • 在训练中动态优化分词策略,支持预训练与微调
  • 三种语言分词演变呈现四阶段规律,复杂语言更不稳定
  • 微调时分词变细粒度,适合低资源复杂语言任务

传统子词分词通常在预处理阶段固定,而本文研究其在训练中的动态学习机制:语言模型能否在训练过程中动态优化分词方式?为此,我们扩展了子词分段语言模型(SSLM)框架,使其支持预训练和微调。以三种形态学差异显著的语言为对象:构词复杂的祖鲁语(Isi-Xhosa)、形态分离的茨瓦纳语(Setswana)及中等形态特征的英语。从语言学角度分析子词演化,追踪形态、产出性与繁殖力。发现子词学习存在四个阶段,形态复杂的祖鲁语表现出更高不稳定性。微调阶段子词边界趋于细化。结果表明,可学习子词为低资源、形态复杂的语言在文本生成与跨语言迁移中提供了有效方案。

原文摘要 · Abstract (English)

Subword segmentation is typically applied in preprocessing and stays fixed during training. Alternatively, it can be learned during training to optimise the training objective. In this paper we study the learning dynamics of subword segmentation: if a language model can dynamically optimise tokenisation, how do its subwords evolve during pretraining and finetuning? To explore this, we extend the subword segmental language model (SSLM), a framework for learning subwords during training, to support pretraining and finetuning. We train models for three typologically diverse languages to study learning dynamics across the morphological spectrum: Isi-Xhosa is conjunctive (long word forms composed of many morphemes), Setswana is disjunctive (morphemes written as separate words), and English represents a typological middle ground. We analyse subword dynamics from a linguistic perspective, tracking morphology, productivity, and fertility. We identify four stages of subword learning, with the morphologically complex isi-Xhosa exhibiting greater instability. During finetuning, subword boundaries shift to become finer-grained. Lastly, we show that learnable subwords offers a promising approach to improve text generation and cross-lingual transfer for low-resource, morphologically complex languages.

子词分词形态学语言模型低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。