提出联合预测分词与释义的多语言模型,提升语言标注可信度。
Massively Multilingual Joint Segmentation and Glossing
- 首次构建端到端联合分词与释义的神经网络模型
- 在分词和释义任务上均优于GlossLM和开源大模型
- 支持快速适配新语料,适合语言学家和田野调查者使用
神经网络自动预测逐行释义是加速语言记录的有前景方法。然而,尽管GlossLM等先进模型在释义基准上表现优异,语言学家的用户研究发现其在实际应用中存在关键障碍:现有模型通常生成词素级释义但不预测真实词素边界,导致结果难以理解且缺乏可信度。本文首次研究联合从原始文本中预测逐行释义与对应形态分割的神经模型。通过实验确定了平衡分词与释义准确率及任务对齐的最优训练方式。我们扩展了GlossLM的训练语料,预训练了PolyGloss系列序列到序列多语言模型,在释义任务上超越GlossLM,且在分词、释义和对齐三项指标上均优于多种开源大模型。此外,我们证明了PolyGloss可通过低秩适配快速适应新数据集。
原文摘要 · Abstract (English)
Automated interlinear gloss prediction with neural networks is a promising approach to accelerate language documentation efforts. However, while state-of-the-art models like GlossLM achieve high scores on glossing benchmarks, user studies with linguists have found critical barriers to the usefulness of such models in real-world scenarios. In particular, existing models typically generate morpheme-level glosses but assign them to whole words without predicting the actual morpheme boundaries, making the predictions less interpretable and thus untrustworthy to human annotators. We conduct the first study on neural models that jointly predict interlinear glosses and the corresponding morphological segmentation from raw text. We run experiments to determine the optimal way to train models that balance segmentation and glossing accuracy, as well as the alignment between the two tasks. We extend the training corpus of GlossLM and pretrain PolyGloss, a family of seq2seq multilingual models for joint segmentation and glossing that outperforms GlossLM on glossing and beats various open-source LLMs on segmentation, glossing, and alignment. In addition, we demonstrate that PolyGloss can be quickly adapted to a new dataset via low-rank adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。