用多智能体系统自动生成符合课标的科学测验题,更准更可靠。
TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

- 分层知识库+多阶段智能体流水线,精准匹配知识点与教学水平
- 生成题目的准确率提升至0.96,答案相关性达0.89,远超传统方法
- 专为低资源课本设计,适合教育工作者和智能出题研究者
自动生成基于教科书的测评题可减轻科学教师负担,但现有检索增强生成(RAG)系统依赖扁平化检索,仅支持单题生成,缺乏对弱证据的防护,且不适用于低资源、以考试为导向的课程体系。我们提出TeachMateGPT,一个面向课程教材的多智能体知识增强型测评生成框架,带来四项改进:(i) COPE——分层知识库,按课程结构分段文档,通过图谱关联三种粒度信息,精准匹配知识点与教学层级;(ii) 分阶段、故障封闭式智能体流水线,取代单次检索生成:路由门搜索,检索融合密集与词法证据,覆盖率不足则阻止生成,专业智能体分别撰写选择题与开放题;(iii) SAVER——带来源标注的验证协议,对每道题四个子部分进行忠实度、相关性与幻觉风险评分,结合教师参与评估而非自动过滤;(iv) NCTB-SciGen8——首个基于印度国家课程委员会八年级科学教材的评测数据集,含198道题(143道选择题,55道开放题),覆盖全部14章,由流水线生成并经三位在职教师评分。TeachMateGPT将忠实度从0.68提升至0.96,答案相关性从0.60提升至0.89,显著优于基础RAG模型。
原文摘要 · Abstract (English)
Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula. We address these limitations with TeachMateGPT, a multi-agent system contributing four advances to curriculum-grounded science-assessment authoring. (i) COPE, a hierarchical knowledge base replacing token-window chunking with a multi-resolution index that segments documents along syllabus structure and links them at three granularities via a traversable graph-based lineage, matching evidence to each topic's instructional level. (ii) A staged, fail-closed agent pipeline replacing one-shot retrieve-then-generate: routing gates search, retrieval fuses dense and lexical evidence under a coverage gate that withholds generation on insufficient evidence, and specialist agents draft objective and constructed-response items. (iii) SAVER, a source-attributed verification protocol scoring faithfulness, relevance, and hallucination risk against retrieved evidence, applying stricter grounding checks across each creative question's four sub-parts, paired with teacher-in-the-loop evaluation rather than automatic filtering. (iv) NCTB-SciGen8, a curriculum-grounded dataset of 198 items (143 multiple-choice, 55 creative questions) spanning all 14 chapters of the NCTB Class 8 science textbook, produced by the pipeline and rated by three practicing teachers. TeachMateGPT raises faithfulness (0.68 $\rightarrow$ 0.96) and answer relevancy (0.60 $\rightarrow$ 0.89) over a vanilla RAG baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。