用概念图指导大模型生成高质量多选题,精准打击常见误解。
Harnessing Structured Knowledge: A Concept Map-Based Approach for High-Quality Multiple Choice Question Generation with Effective Distractors
- 基于分层概念图提供结构化知识,引导大模型生成题目与干扰项。
- 专家评估显示75.2%的题目达标,学生猜对率降至28.05%。
- 适合教育科技、智能测评系统研发人员快速构建高阶测验。
生成高质量多选题,尤其是覆盖不同认知层次并融入常见误解的干扰项,耗时且依赖专业经验,难以规模化。现有自动化方法多生成低阶问题,无法融入领域特定误解。本文提出一种基于分层概念图的框架,为大模型生成题目和干扰项提供结构化知识支持。以高中物理为测试领域,构建涵盖主要知识点及其关联关系的分层概念图,并设计高效数据库。通过自动化流程检索相关概念图片段,作为大模型生成题目的结构化上下文,精准针对常见误解。最后进行自动化验证,确保题目满足预设标准。在两个基线方法(基础大模型与RAG生成)上进行对比评估,包含专家评价与学生测试。专家评估显示,本方法达标率为75.20%,显著优于基线的约37%;学生测试表明,本方法猜对率仅为28.05%,低于基线的37.10%,说明更有效检验概念理解。结果证明,该方法可在多认知层级实现可靠评估,实时识别概念盲区,支持快速反馈与精准干预。
原文摘要 · Abstract (English)
Generating high-quality MCQs, especially those targeting diverse cognitive levels and incorporating common misconceptions into distractor design, is time-consuming and expertise-intensive, making manual creation impractical at scale. Current automated approaches typically generate questions at lower cognitive levels and fail to incorporate domain-specific misconceptions. This paper presents a hierarchical concept map-based framework that provides structured knowledge to guide LLMs in generating MCQs with distractors. We chose high-school physics as our test domain and began by developing a hierarchical concept map covering major Physics topics and their interconnections with an efficient database design. Next, through an automated pipeline, topic-relevant sections of these concept maps are retrieved to serve as a structured context for the LLM to generate questions and distractors that specifically target common misconceptions. Lastly, an automated validation is completed to ensure that the generated MCQs meet the requirements provided. We evaluate our framework against two baseline approaches: a base LLM and a RAG-based generation. We conducted expert evaluations and student assessments of the generated MCQs. Expert evaluation shows that our method significantly outperforms the baseline approaches, achieving a success rate of 75.20% in meeting all quality criteria compared to approximately 37% for both baseline methods. Student assessment data reveal that our concept map-driven approach achieved a significantly lower guess success rate of 28.05% compared to 37.10% for the baselines, indicating a more effective assessment of conceptual understanding. The results demonstrate that our concept map-based approach enables robust assessment across cognitive levels and instant identification of conceptual gaps, facilitating faster feedback loops and targeted interventions at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。