arXiv:2510.07074cs.CLcs.AI2025-10中稿 · KONVENS 2026被引 1

为卢森堡语构建跨语言指令数据集,提升低资源语言模型性能。

LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish

  • 利用英法德三语对齐数据构建卢森堡语指令数据,避免机器翻译误差。
  • 跨语言指令微调显著提升卢森堡语生成能力与多语言表征对齐。
  • 适合关注低资源语言、跨语言模型与数据质量的研究者。

指令微调已成为提升大语言模型性能的关键技术,使其更准确响应人类提示。然而,卢森堡语等低资源语言因缺乏高质量指令数据而受限,传统依赖机器翻译常引入语义偏差和文化失真。本文提出一种无需机器翻译的跨语言指令微调数据集构建方法,基于英语、法语和德语的对齐数据,构建高质量卢森堡语数据集,保留语言与文化细节。实验证明,跨语言指令微调不仅能增强多语言间表征对齐,还能显著提升模型在卢森堡语中的生成能力,表明优质跨语言数据构建可规避机器翻译的常见问题,直接推动低资源语言发展。

原文摘要 · Abstract (English)

Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages such as Luxembourgish face severe limitations due to the lack of high-quality instruction datasets. Traditional reliance on machine translation often introduces semantic misalignment and cultural inaccuracies. In this work, we address these challenges by creating a cross-lingual instruction tuning dataset for Luxembourgish, without resorting to machine-generated translations into it. Instead, by leveraging aligned data from English, French, and German, we build a high-quality dataset that preserves linguistic and cultural nuances. We provide evidence that cross-lingual instruction tuning not only improves representational alignment across languages but also the model's generative capabilities in Luxembourgish. This highlights how cross-lingual data curation can avoid the common pitfalls of machine-translated data and directly benefit low-resource language development.

指令微调低资源语言跨语言数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。