用合成数据提升卢森堡语大模型能力,效果显著。
LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
- 从本地文本生成指令数据,用大模型自检质量
- 14个模型在语言考试中平均准确率提升5.37个百分点
- 适合低资源语言研究者和多语言AI开发者
低资源语言因高质量训练数据匮乏,导致指令微调大模型效果受限。本文提出卢森堡语指令微调数据集LuxIT,基于本土卢森堡语文本,利用DeepSeek-R1-0528模型生成内容,经大模型评审后保留227,507条高质量指令-回答对。为验证实用性,我们在14个参数量≤150亿的较小模型上使用LuxIT进行微调,并在标准化卢森堡语水平测试及五个下游NLP任务上评估。结果显示,所有14个模型在语言考试中平均准确率提升5.37个百分点,其中12个模型表现改善;在9个模型的宏平均F1指标上有所提升,但两个基准任务上的增益无系统性关联。结果表明,利用单语合成数据可有效提升低资源语言的大模型能力,同时揭示语言能力的多维度特性。
原文摘要 · Abstract (English)
The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quality training data. We introduce LuxIT, a monolingual instruction tuning dataset for Luxembourgish developed to mitigate this challenge. We synthesize the dataset from a corpus of native Luxembourgish texts, utilizing DeepSeek-R1-0528, chosen for its shown proficiency in Luxembourgish. Following generation, we apply a quality assurance process, employing an LLM-as-a-judge approach, retaining 227,507 high-quality instruction-answer pairs. To investigate the practical utility of the dataset, we fine-tune 14 smaller-scale LLMs ($\leq$15B parameters) on LuxIT and evaluate them on standardized Luxembourgish proficiency exams and five downstream NLP tasks. Training on LuxIT yields a mean accuracy change of +5.37 percentage points on language exams across all 14 models, with 12 of 14 showing improvement. On NLP downstream tasks, 9 of 14 models improve in macro-averaged F1, though gains on the two benchmarks do not systematically correlate. These results underscore the feasibility of leveraging monolingual synthetic data to improve LLM capabilities in low-resource languages, while highlighting the multi-faceted nature of language proficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。