用生成模型为濒危语言创建数据,助力语言复兴
LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models
- 基于生成模型自动生成低资源语言数据
- 以撒丁语为例验证框架有效性
- 适合语言保护与技术复兴研究者使用
少数语言对文化遗产的保存至关重要,但因数字资源匮乏及高资源语言模型的主导地位,正面临日益严重的灭绝风险。本文提出一个开源框架,利用生成模型为低资源语言生成语言工具,重点解决数据创建问题,以支持语言模型开发,助力语言保护。以濒危语言撒丁语为案例,验证了该框架在缓解数据稀缺方面的有效性。本工作推动语言多样性发展,通过现代技术手段支持语言标准化与复兴努力。
原文摘要 · Abstract (English)
Minority languages are vital to preserving cultural heritage, yet they face growing risks of extinction due to limited digital resources and the dominance of artificial intelligence models trained on high-resource languages. This white paper proposes a framework to generate linguistic tools for low-resource languages, focusing on data creation to support the development of language models that can aid in preservation efforts. Sardinian, an endangered language, serves as the case study to demonstrate the framework's effectiveness. By addressing the data scarcity that hinders intelligent applications for such languages, we contribute to promoting linguistic diversity and support ongoing efforts in language standardization and revitalization through modern technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。