arXiv:2605.09635cs.CL2026-05

构建首个面向K12教育的知识图谱,用于评估和训练具备课程认知能力的模型。

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

论文配图:K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
图 1 · 摘自论文原文
  • 从人教版教材提取九类节点十四类关系,构建课程对齐知识图谱。
  • 创建2.36万题多选基准,其中前置条件与邻近概念任务最难,准确率不足57%。
  • 融合文本与视觉监督的训练数据表现最优,适合教育大模型训练与评测。

大语言模型在K-12教育中应用日益广泛,但现有基准多聚焦于考试问答,缺乏对课程知识结构与可视化呈现的理解能力评估,即课程认知能力,涵盖前提链、概念层级、实验-概念关联、教学顺序及视觉定位。本文提出K12-KGraph,基于人民教育出版社数学、物理、化学、生物教材构建的课程对齐知识图谱,包含九种节点类型与十四种关系类型,覆盖课程结构与视觉定位。由此衍生出包含23,640道题的多选基准K12-Bench,涵盖五类任务:定位、前提、邻近、证据与发现。同时构建含7,335样本的图引导微调语料库K12-Train,包括2,267个纯文本问答对与5,068个多模态视觉问答对。在K12-Bench上,Gemini-3-Flash仅达57%精确匹配,Gemma-4-31B-IT为46%,其中前置条件与邻近任务最困难。训练实验表明,领域特定监督可缩小差距。在2,300样本预算下,K12-Train-Text持续优于八个主流指令微调语料库子集,在GaokaoBench与EduEval上表现更优。对于视觉语言模型,尽管样本少于DataFlow与WizardLM基线,K12-Train-Full在Gaokao-MM、MDK12-medium与K12Vista上均取得最佳结果,且超越纯文本与纯视觉变体,证明文本与视觉监督具有互补性。相关数据集、基准、训练数据及构建流程已开源。

原文摘要 · Abstract (English)

Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbooks in mathematics, physics, chemistry, and biology across primary, middle, and high school. It contains nine node types and fourteen relation types covering curriculum structure and visual grounding. From this graph, we derive K12-Bench, a 23,640-question multi-select benchmark with five task families: Ground, Prereq, Neighbor, Evidence, and Locate. We also build K12-Train, a graph-guided supervised fine-tuning corpus of 7,335 samples, including 2,267 text-only QA pairs and 5,068 multimodal VQA pairs. On K12-Bench, Gemini-3-Flash achieves only 57 percent exact match and Gemma-4-31B-IT reaches 46 percent, with Prereq and Neighbor being the hardest tasks. Our training experiments show that domain-specific supervision can reduce this gap. Under a matched 2,300-sample budget, K12-Train-Text consistently outperforms equally sized subsets of eight mainstream instruction-tuning corpora on GaokaoBench and EduEval. For vision-language models, K12-Train-Full achieves the best overall results on Gaokao-MM, MDK12-medium, and K12Vista among all compared training configurations, despite using fewer samples than the full DataFlow and WizardLM baselines. It also surpasses both text-only and multimodal-only variants, showing that textual and visual supervision are complementary. We release the graph, benchmark, training data, and complete construction pipeline.

教育AI知识图谱多模态大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。