用大模型生成符合印度课纲的互动教材,提升学习个性化体验。
PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
- 基于课程内容构建分级问答数据集,分类为事实、推理与综合题。
- 对比多种提示策略,发现思维链提示在推理题上表现最优。
- 评估开源与闭源模型,适合教育场景的中小模型已具实用价值。
大型语言模型(LLMs)在理解与生成类人内容方面展现出卓越能力,已深刻影响医疗、软件开发与教育等领域。在教育中,它们可提供个性化、互动式学习体验,尤其适用于教学资源匮乏地区。然而,将这些模型有效适配至特定课程内容(如印度国家教育研究委员会NCERT课纲)仍面临准确性、内容对齐性与教学相关性的挑战。本文提出PustakAI框架,用于设计与评估一个新型问答数据集NCERT-QA,覆盖6至8年级英语与科学科目。我们对收集的问答对进行分类:事实型、推理性及其它(评价与推理类)。通过元提示、少样本提示和思维链提示等多种方法,在多维度评估指标下测试模型表现,探究何种策略更契合课程结构需求。同时分析当前开源模型(Gemma3:1b、Llama3.2:3b、Nemotron-mini:4b)与高端模型(Llama-4-Scout-17B、Deepseek-r1-70B)作为教育工具的优劣,验证其在正式教育系统中的可用性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generating human-like content. This has revolutionized various sectors such as healthcare, software development, and education. In education, LLMs offer potential for personalized and interactive learning experiences, especially in regions with limited teaching resources. However, adapting these models effectively to curriculum-specific content, such as the National Council of Educational Research and Training (NCERT) syllabus in India, presents unique challenges in terms of accuracy, alignment, and pedagogical relevance. In this paper, we present the framework "PustakAI"\footnote{Pustak means `book' in many Indian languages.} for the design and evaluation of a novel question-answering dataset "NCERT-QA" aligned with the NCERT curriculum for English and Science subjects of grades 6 to 8. We classify the curated QA pairs as Factoid, Inferential, and Others (evaluative and reasoning). We evaluate the dataset with various prompting techniques, such as meta-prompt, few-shot, and CoT-style prompting, using diverse evaluation metrics to understand which approach aligns more efficiently with the structure and demands of the curriculum. Along with the usability of the dataset, we analyze the strengths and limitations of current open-source LLMs (Gemma3:1b, Llama3.2:3b, and Nemotron-mini:4b) and high-end LLMs (Llama-4-Scout-17B and Deepseek-r1-70B) as AI-based learning tools in formal education systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。