用GPT-4o自动生成阅读理解推断题,提升诊断评估效率。
Automatic Generation of Inference Making Questions for Reading Comprehension Assessments
- 基于推理类型分类,用少样本提示生成推断题。
- 93.8%生成题目质量达标,但仅42.6%匹配目标推理类型。
- 适合教育机构批量构建诊断性阅读测试,需人工校验。
推断能力是阅读理解的核心但复杂技能,部分推断需跨句指代消解,部分依赖背景知识填补文本未明说内容。诊断性阅读题可帮助教师为学龄学生提供更精准的教学与干预。本文提出阅读理解推理类型的分类体系,并分析诊断题库中题目分布。接着,利用GPT-4o通过少样本提示生成桥接推理题,对比有无思维链提示的条件。生成题目在整体质量、推理类型匹配度和大模型推理合理性上均达到0.90以上的评分者间一致性。结果表明,GPT-4o生成的93.8%题目质量良好,适用于3-12年级场景;但仅有42.6%准确匹配目标推理类型。结论:自动出题结合人工判断,是实现规模化高质量诊断性阅读评估的可行路径。
原文摘要 · Abstract (English)
Inference making is an essential but complex skill in reading comprehension (RC). Some inferences require resolving references across sentences, and some rely on using prior knowledge to fill in the detail that is not explicitly written in the text. Diagnostic RC questions can help educators provide more effective and targeted reading instruction and interventions for school-age students. We introduce a taxonomy of inference types for RC and use it to analyze the distribution of items within a diagnostic RC item bank. Next, we present experiments using GPT-4o to generate bridging-inference RC items for given reading passages via few-shot prompting, comparing conditions with and without chain-of-thought prompts. Generated items were evaluated on three aspects: overall item quality, appropriate inference type, and LLM reasoning, achieving high inter-rater agreements above 0.90. Our results show that GPT-4o produced 93.8% good-quality questions suitable for operational use in grade 3-12 contexts; however, only 42.6% of the generated questions accurately matched the targeted inference type. We conclude that combining automatic item generation with human judgment offers a promising path toward scalable, high-quality diagnostic RC assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。