arXiv:2511.17161cs.CLcs.AI2025-11被引 2

构建波兰语大模型的指令数据集,区分人工、转换与合成指令类型。

The PLLuM Instruction Corpus

  • 按来源分类:人工、转换、合成指令,形成完整数据类型体系
  • 发布首个代表性的波兰语指令子集PLLuMIC,供后续研究使用
  • 对比人工与合成数据对模型语言适应的影响,提供实证参考

本文介绍了用于微调波兰语大语言模型(LLMs)的指令数据集,该模型基于Transformer架构,出自PLLuM(波兰大语言模型)项目。我们提出了一种功能性分类体系,涵盖人工生成、转换而来和合成的指令,并分享了关于在基础大模型语言适配中使用人工与合成指令数据集影响的观察。此外,我们发布了首个具有代表性的PLLuM指令语料库子集(PLLuMIC),相信其对其他大模型类似数据集的开发具有指导意义。

原文摘要 · Abstract (English)

This paper describes the instruction dataset used to fine-tune a set of transformer-based large language models (LLMs) developed in the PLLuM (Polish Large Language Model) project. We present a functional typology of the organic, converted, and synthetic instructions used in PLLuM and share some observations about the implications of using human-authored versus synthetic instruction datasets in the linguistic adaptation of base LLMs. Additionally, we release the first representative subset of the PLLuM instruction corpus (PLLuMIC), which we believe to be useful in guiding and planning the development of similar datasets for other LLMs.

指令数据集波兰语大模型语言适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。