为濒危语言拉丁语构建首个公开的文本分析数据集
Exploring NLP Benchmarks in an Extremely Low-Resource Setting
- 用意语语料合成拉丁语情感分析和问答数据
- 合成数据使机器翻译性能显著优于现有基线
- 适合关注小语种NLP与濒危语言保护的研究者
大语言模型在极低资源语言(如土著语言)上的表现下降,主要因标注数据匮乏。尽管关注度上升,但高质量自然语言处理数据集仍极为有限,难以支撑可靠语言技术开发。本文聚焦于濒危罗曼语族语言拉丁语的瓦尔巴迪亚变体,利用少量平行的拉丁语-意大利语句子对,通过翻译单语意大利语语料生成情感分析和多选问答(MCQA)的合成数据集。为保证语言质量与可靠性,方法中采用严格筛选与回译流程。实验表明,将合成数据融入机器翻译训练后,其性能显著优于现有意大利语-拉丁语翻译基线。本研究贡献了首个公开的拉丁语情感分析与MCQA数据集,为该代表性不足语言的NLP研究及下游应用奠定基础。
原文摘要 · Abstract (English)
The effectiveness of Large Language Models (LLMs) diminishes for extremely low-resource languages, such as indigenous languages, primarily due to the lack of labeled data. Despite growing interest, the availability of high-quality natural language processing (NLP) datasets for these languages remains limited, making it difficult to develop robust language technologies. This paper addresses such gap by focusing on Ladin, an endangered Romance language, specifically targeting the Val Badia variant. Leveraging a small set of parallel Ladin-Italian sentence pairs, we create synthetic datasets for sentiment analysis and multiple-choice question answering (MCQA) by translating monolingual Italian data. To ensure linguistic quality and reliability, we apply rigorous filtering and back-translation procedures in our method. We further demonstrate that incorporating these synthetic datasets into machine translation training leads to substantial improvements over existing Italian-Ladin translation baselines. Our contributions include the first publicly available sentiment analysis and MCQA datasets for Ladin, establishing foundational resources that can support broader NLP research and downstream applications for this underrepresented language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。