用AI生成符合教学目标的编程选择题,人机协作提升题目质量。
CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
- AI分两步:生成题干,再独立验证七项教育维度。
- 专家评估显示准确率79.9%至98.6%,代码与概念对齐度高。
- 适合需要高效出题的教育开发者,尤其擅长可计算验证项。
我们提出CODE-GEN,一个基于检索增强生成(RAG)的人机协同智能体系统,用于生成与课程学习目标对齐的编程理解选择题,以提升学生代码推理与理解能力。CODE-GEN采用智能体架构:生成器(Generator)根据课程目标生成选择题,验证器(Validator)独立评估七个教育维度的内容质量。两个智能体均配备专用工具,提升计算准确性并验证代码输出。为评估有效性,六位领域专家(SMEs)对288道AI生成题目进行了评判,产生共2016组人机评分对比及131条定性反馈。分析显示,人类验证成功率在七个维度间为79.9%至98.6%。定性反馈表明,系统在问题清晰度、代码正确性、概念一致性及正确答案有效性等可计算维度表现优异;而在设计有意义干扰项与提供高质量反馈等需深度教学判断的维度,仍需人类专家介入。研究结果为教育内容生成中人机任务分工提供了依据。
原文摘要 · Abstract (English)
We present CODE-GEN, a human-in-the-Loop, retrieval-augmented generation (RAG)-based agentic AI system for generating context-aligned multiple-choice questions to develop student code reasoning and comprehension abilities. CODE-GEN employs an agentic AI architecture in which a Generator agent produces multiple-choice coding comprehension questions aligned with course-specific learning objectives, while a Validator agent independently assesses content quality across seven pedagogical dimensions. Both agents are augmented with specialized tools that enhance computational accuracy and verify code outputs. To evaluate the effectiveness of CODE-GEN, we conducted an evaluation study involving six human subject-matter experts (SMEs) who judged 288 AI-generated questions. The SMEs produced a total of 2,016 human-AI rating pairs, indicating agreement or disagreement with the assessments of Validator, along with 131 instances of qualitative feedback. Analyses of SME judgments show strong system performance, with human-validated success rates ranging from 79.9% to 98.6% across the seven pedagogical dimensions. The analysis of qualitative feedback reveals that CODE-GEN achieves high reliability on dimensions well suited to computational verification and explicit criteria matching, including question clarity, code validity, concept alignment, and correct answer validity. In contrast, human expertise remains essential for dimensions requiring deeper instructional judgment, such as designing pedagogically meaningful distractors and providing high-quality feedback that reinforces understanding. These findings inform the strategic allocation of human and AI effort in AI-assisted educational content generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。