用大模型自动生成藏语思维链数据,解决藏语低资源难题
TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- 用大模型思维链提示自动构建藏语多领域数据集
- 自建的Sunshine-thinking模型在藏语推理上达SOTA水平
- 适合关注少数民族语言AI、低资源语言处理的研究者
为解决藏语这一超六百万人口使用的低资源语言严重缺乏数据的问题,我们提出TIBSTC-CoT——一个通过大语言模型思维链提示自动生成的大规模、多领域藏语数据集。该数据集构建了可扩展、可复现的低资源语言数据创建框架,覆盖多样领域与推理模式,对语言理解与生成至关重要。基于此数据集,我们开发了Sunshine-thinking系列藏语专用大模型,具备思维链能力。全部在TIBSTC-CoT上训练,其推理与生成性能媲美顶尖多语言大模型。本工作推动包容性AI发展,实现高质量藏语语言处理。所有数据已公开:https://github.com/Vicentvankor/sun-shine。
原文摘要 · Abstract (English)
To address the severe data scarcity in Tibetan, a low-resource language spoken by over six million people, we introduce TIBSTC-CoT, the large-scale, multi-domain Tibetan dataset automatically constructed via chain-of-thought prompting with large language models (LLMs). TIBSTC-CoT establishes a scalable and reproducible framework for dataset creation in low-resource settings, covering diverse domains and reasoning patterns essential for language understanding and generation. Building on this dataset, we develop the Sunshine-thinking LLM family, a series of Tibetan-centric LLMs equipped with chain-of-thought capabilities. Trained entirely on TIBSTC-CoT, Sunshine-thinking has demonstrated strong reasoning and generation performance, comparable to state-of-the-art (SOTA) multilingual LLMs. Our work marks a significant step toward inclusive AI by enabling high-quality Tibetan language processing through both resource creation and model innovation. All data are available: https://github.com/Vicentvankor/sun-shine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。