为三种濒危芬兰-乌戈尔语开发了大语言模型,推动低资源语言的NLP发展。
LLMs for Extremely Low-Resource Finno-Ugric Languages
- 针对三种低资源语言构建多语言基础与指令微调模型
- 创建了smugri-MT-bench多轮对话评估基准,支持人类评测
- 覆盖从数据收集到评估的完整LLM研发流程,助力语言多样性
大型语言模型(LLMs)的发展主要集中在高资源语言,导致如芬兰-乌戈尔语系中的维罗语、利沃尼亚语和科米语等低资源语言严重缺乏支持。本文聚焦这三种语言,覆盖了从数据收集到指令微调及评估的完整LLM研发周期。贡献包括构建多语言基础模型与指令微调模型;开发评估基准,包含smugri-MT-bench多轮对话评测集;并开展人工评估。本工作旨在促进语言多样性,确保低资源语言也能从自然语言处理技术进步中受益。
原文摘要 · Abstract (English)
The advancement of large language models (LLMs) has predominantly focused on high-resource languages, leaving low-resource languages, such as those in the Finno-Ugric family, significantly underrepresented. This paper addresses this gap by focusing on Võro, Livonian, and Komi. We cover almost the entire cycle of LLM creation, from data collection to instruction tuning and evaluation. Our contributions include developing multilingual base and instruction-tuned models; creating evaluation benchmarks, including the smugri-MT-bench multi-turn conversational benchmark; and conducting human evaluation. We intend for this work to promote linguistic diversity, ensuring that lesser-resourced languages can benefit from advancements in NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。