arXiv:2411.02556cs.CL2024-11被引 2

用Transformer模型预测濒危萨米语的词形变化类别,助力语言保护。

Leveraging Transformer-Based Models for Predicting Inflection Classes of Words in an Endangered Sami Language

  • 基于Transformer构建端到端模型,处理数据稀缺与复杂形态
  • 词性分类F1达1.00,词形变化类别分类达0.81
  • 成果开源,适合语言保护与小语种NLP研究者使用

本文提出一种训练Transformer模型的方法,用于分类斯科特萨米语(Skolt Sami)的词汇与形态句法特征。该语言属乌拉尔语系,形态复杂且数据稀缺。研究构建了从数据提取、增强到模型训练的全流程系统,旨在提升对斯科特萨米语的理解与分析能力。准确分类不仅可增强有限状态转换器(FST)的词库覆盖,还支持研究人员对文献与母语者新发现词汇的系统性记录。模型在词性分类上取得平均加权F1分数1.00,在词形变化类别分类上为0.81。相关模型与代码将公开发布,以促进濒危语言自然语言处理研究。

原文摘要 · Abstract (English)

This paper presents a methodology for training a transformer-based model to classify lexical and morphosyntactic features of Skolt Sami, an endangered Uralic language characterized by complex morphology. The goal of our approach is to create an effective system for understanding and analyzing Skolt Sami, given the limited data availability and linguistic intricacies inherent to the language. Our end-to-end pipeline includes data extraction, augmentation, and training a transformer-based model capable of predicting inflection classes. The motivation behind this work is to support language preservation and revitalization efforts for minority languages like Skolt Sami. Accurate classification not only helps improve the state of Finite-State Transducers (FSTs) by providing greater lexical coverage but also contributes to systematic linguistic documentation for researchers working with newly discovered words from literature and native speakers. Our model achieves an average weighted F1 score of 1.00 for POS classification and 0.81 for inflection class classification. The trained model and code will be released publicly to facilitate future research in endangered NLP.

濒危语言Transformer形态学语言保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。