用知识图谱和优化算法挖掘考试隐藏能力结构,实现个性化备考计划
LearnOpt: Recovering the Latent Cognitive Structure of Standardized Examinations via Knowledge Graphs and Constrained Optimization
- 基于历史真题构建知识图谱,通过约束优化提取五类隐含能力分布
- 发现考试能力结构在课程调整前稳定,2023年教材改革后显著变化(KL=0.040)
- 适用于备考策略研究者、教育AI开发者及关注考试本质的教师
标准化考试常被视为统一的知识覆盖问题,但我们认为其本质是具有稳定隐含认知结构的对抗性系统,且该结构与官方大纲系统性偏离。本文提出LearnOpt,从历年真题中恢复这一结构,并生成个性化的限时学习计划。以2016-2024年共1496道NEET真题为样本,利用LLM标注构建考试知识图谱,提取出五类隐含技能分布,并将学习规划建模为带先决条件的子图优化问题,结合贝叶斯知识追踪。核心发现:在2016-2021年课程体系下,隐含技能分布保持稳定(连续年份间KL散度为0.004-0.032,置换检验不显著);而2023年NCERT课程改革后,技能分布发生显著变化(2016-2021年n=1072 vs 2023-2024年n=392,KL=0.040,p=0.0005),其中排除法/否定类题目占比从约20%-29%升至约31%-35%。尽管非永久不变,但结构呈现分段稳定特征,且可归因于课程变动。在同一课程阶段内,科目比年份更能预测技能分布。优化评估显示,基于技能权重的目标相较基础频率基准,能带来明显但温和的学习建议重排序。将该方法应用于JEE Advanced,发现其技能分布以多概念整合为主(80.9%对比NEET的33.3%),且JEE与NEET间的差异(KL=0.505)超过NEET内部最大跨科目差异,表明考试层级对隐含认知结构的影响大于科目,科目又大于时间因素。代码、知识图谱与标注数据集已公开。
原文摘要 · Abstract (English)
Standardized examinations are typically treated as uniform syllabus coverage problems. We argue they are better understood as adversarial systems with stable latent cognitive structures diverging systematically from official syllabi. We introduce LearnOpt, which recovers this structure from historical question papers and generates personalized, time-bounded study plans. Applied to nine years of NEET questions (2016-2024, n=1,496), LearnOpt builds an exam knowledge graph from LLM-tagged questions, extracts a five-category latent skill distribution, and formulates study planning as a knapsack-variant optimization over prerequisite-aware subgraphs with Bayesian Knowledge Tracing. Central finding: NEET's latent skill distribution is stable within a syllabus regime (consecutive-year KL divergence 0.004-0.032 for 2016-2021, non-significant under permutation testing) but shifts significantly with NCERT's 2023 syllabus rationalization: pooling 2016-2021 (n=1,072) vs 2023-2024 (n=392) gives KL=0.040 (p=0.0005), with Elimination/Negation questions rising from ~20-29% to ~31-35%. Latent structure, while not permanently stationary, is piecewise stable, with shifts detectable and attributable to curricular events. Within either regime, subject predicts skill profile more strongly than year. An optimization evaluation, using one real and two synthetic mastery profiles, shows the skill-weighted objective produces a modest but real reordering of recommended topics over a mastery-conditioned frequency baseline. Applying the pipeline to JEE Advanced reveals a profile dominated by Multi-concept Integration (80.9% vs. 33.3% for NEET), with a JEE-vs-NEET divergence (KL=0.505) exceeding NEET's largest cross-subject divergence: exam tier shapes latent cognitive structure more than subject, which shapes it more than time within a regime. Code, knowledge graph, and annotated dataset are released publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。