用多任务学习和大模型生成数据,提升低资源语言的词素分割准确率。
Learning Beyond Limits: Multitask Learning and Synthetic Data for Low-Resource Canonical Morpheme Segmentation
- 通过共享语言表示,联合预测词素与释义,增强泛化能力。
- 在多个低资源语言上,词级分割准确率和词素级F1值均显著提升。
- 适合需要少样本词素分析的研究者,尤其关注低资源语言处理。
我们提出一种基于Transformer的词素分割系统,通过多任务学习和大语言模型(LLM)生成的合成数据来增强低资源训练信号。该框架联合从正字法输入中预测形态学片段与释义,利用共同文档处理过程获得的共享语言表征以提升模型泛化能力。为应对数据稀缺问题,我们引入通过上下文学习生成的合成训练数据。在SIGMORPHON 2023数据集上的实验表明,该方法在多个低资源语言上显著提升了词级分割准确率和词素级F1分数。
原文摘要 · Abstract (English)
We introduce a transformer-based morpheme segmentation system that augments a low-resource training signal through multitask learning and LLM-generated synthetic data. Our framework jointly predicts morphological segments and glosses from orthographic input, leveraging shared linguistic representations obtained through a common documentary process to enhance model generalization. To further address data scarcity, we integrate synthetic training data generated by large language models (LLMs) using in-context learning. Experimental results on the SIGMORPHON 2023 dataset show that our approach significantly improves word-level segmentation accuracy and morpheme-level F1-score across multiple low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。