arXiv:2607.23322cs.CLcs.CY2026-07

构建2.4万条多语言印度知识教学数据,让大模型学会传授本土教育内容。

IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems

论文配图:IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems
图 1 · 摘自论文原文
  • 从古籍与课程中提炼41种印度教学法,覆盖7种语言
  • 70亿参数模型微调后评分达6.39,接近主流模型表现
  • 适合想推广本土知识体系的教育科技与研究者

指令微调已成为使大语言模型遵循人类意图的标准方法,但现有指令数据集以英语通用知识任务为主,缺乏对专门教学领域的覆盖。本文提出IKS-Instruct,一个包含24,795个指令-回复对的数据集,用于训练语言模型传授基于印度知识体系(IKS)的教育内容。数据集涵盖英语、印地语、梵语、泰米尔语、泰卢固语、卡纳达语和马拉雅拉姆语七种语言,覆盖41种源自吠陀口传与数学传统的教学方法,并与中学6至12年级的中央教育部(CBSE)课程对齐。数据来源包括经典文本语料库(《薄伽梵歌》、《蒂鲁克库拉尔》、桑伽姆文学、吠陀文本)、课程对齐的教学模板、吠陀数学术法演示、双语指令对、基于教学法的多轮对话以及跨传统比较分析。通过多评委评估框架,在1,201个分层样本上由独立语言模型从12个维度(如教学法契合度、教学品质、事实准确性、文化深度)评分。在统一五评委外部评审组(中位数聚合)下,经强化的70亿参数模型微调后,中位评分为6.39,仅比强基准模型Nemotron-Nano(6.54)低0.15,且部署成本极低;未微调的基础模型在IKS相关维度得分接近零。模型质量并未随数据精修单调提升,该现象连同对应的质量提升结果一并报告。

原文摘要 · Abstract (English)

Instruction tuning has become the standard method for adapting large language models to follow human intent, yet existing instruction datasets are dominated by English-language general-knowledge tasks and lack coverage of specialized pedagogical domains. This paper presents IKS-Instruct, a dataset of 24,795 instruction-response pairs for teaching language models to deliver educational content grounded in Indian Knowledge Systems (IKS). The dataset spans seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam), covers 41 pedagogical techniques from the Vedic oral and mathematical traditions, and is aligned with the Central Board of Secondary Education (CBSE) curriculum for classes 6 through 12. The pairs are derived from six source types: classical text corpora (Bhagavad Gita, Thirukkural, Sangam literature, Vedic texts), curriculum-aligned pedagogical templates, Vedic mathematical sutra demonstrations, bilingual instruction pairs, technique-grounded multi-turn dialogues, and cross-tradition comparative analyses. Quality is assessed through a multi-judge evaluation framework in which independent language models score responses on 12 dimensions including technique fidelity, pedagogical quality, factual accuracy, and IKS cultural depth. Under a uniform five-judge external panel (median aggregation over 1,201 stratified items), the strongest IKS-Instruct fine-tune of a compact 7B model reaches a median judge score of 6.39, within 0.15 of a strong general-purpose reference model (Nemotron-Nano at 6.54) at a fraction of its deployment cost, while the base model without IKS fine-tuning scores near zero on the IKS-specific dimensions. Model quality does not increase monotonically with data curation, a result we report together with the corresponding data-quality gains.

多语言教育模型知识体系数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。