arXiv:2503.18288cs.CL2025-03被引 5

首个覆盖藏语大模型全链路训练的结构化数据集,解决低资源语言建模核心瓶颈。

TFD: A Comprehensive Structured Tibetan Foundation Dataset for Low-Resource Language Processing and Large-Scale Modeling

  • 构建涵盖预训练到推理监督的完整数据链条,统一110亿词表
  • 训练出Sun-Shine系列模型,在理解、安全、推理等任务上显著优于基线
  • 提供首个大规模藏语思维链数据集,助力文化适配型AI发展

大型语言模型在高资源语言中取得显著进展,但藏语领域仍受制于数据匮乏。现有工作虽开始缓解预训练数据短缺,但更根本的问题在于:缺乏支持大模型开发全流程(包括预训练、指令微调、安全对齐、偏好优化和推理监督)的系统性资源。本文提出藏语基础数据集TFD,是首个覆盖藏语大模型全阶段的结构化、大规模、专家标注数据集。TFD包含TIBSTC——一个超110亿词元的统一语料库,并细分出用于指令微调、安全对齐和偏好优化的子数据集;以及TIBSTC-CoT——首个大规模藏语思维链数据集。通过训练Sun-Shine系列藏语大模型,我们在理解、安全、推理与生成等多个基准测试中实现显著提升。结果表明,推动低资源语言建模不仅需要数据规模,更需完整的数据生态体系。我们已公开发布TFD,以支持可复现研究及文化契合的藏语大模型开发。代码与数据见 https://github.com/Vicentvankor/sun-shine。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success in high-resource languages, yet progress in Tibetan remains severely constrained. While recent efforts have begun to address pre-training data scarcity for Tibetan, a more fundamental gap persists: no existing resource supports the complete LLM development pipeline, spanning pre-training, instruction tuning, safety alignment, preference optimization, and reasoning supervision. We introduce the Tibetan Foundation Dataset (TFD), the first structured, large-scale, and expert-curated dataset covering all key stages of Tibetan large language modeling. TFD comprises TIBSTC, a unified corpus of over 11 billion tokens with curated sub-datasets for instruction tuning, safety alignment, and preference optimization, and TIBSTC-CoT, the first large-scale Tibetan chain-of-thought dataset. We demonstrate its utility by training the Sun-Shine family of Tibetan LLMs, achieving substantial improvements over strong baselines on understanding, safety, reasoning, and generation benchmarks. These results underscore that advancing low-resource language modeling requires not only scale, but a structurally complete data ecosystem. We release TFD to facilitate reproducible research and the development of robust, culturally aligned Tibetan LLMs. Code and data are available at https://github.com/Vicentvankor/sun-shine.

藏语AI低资源语言大模型数据集思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。