arXiv:2508.08314cs.CYcs.AI2025-08AAAI被引 12

AI可生成媲美专家的高质量考试题,助力教学提效。

Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study

  • 通过迭代式生成与优化,让AI不断自我改进题目质量。
  • 在1700名学生中,AI题与专家题表现相当,信度达0.82。
  • 适合教育科技、教师工具开发及大规模测评场景使用。

尽管大语言模型(LLMs)对传统教与学方式构成挑战,但也为提升教学效率和规模化高质量教学提供了新机遇。一个有前景的应用是生成针对特定课程内容的定制化考试题。虽然近年来人工智能自动生成题目备受关注,但对其在真实教育环境中试题心理测量质量的评估仍相对不足。填补这一空白是理解生成式AI在有效测验设计中作用的关键一步。本研究引入并评估了一种迭代优化策略:通过多轮循环,由大语言模型生成、评估并改进题目,结合其自身批判性反馈进行修订。我们在涵盖计算机科学、数学、化学等领域的91个课堂中开展了大规模实地研究,涉及美国数十所高校的近1700名学生。基于项目反应理论(IRT)的分析表明,对于本样本中的学生,AI生成题目在区分度和信度上与专家设计的标准化考试题表现相当,整体信度为0.82。结果表明,AI有能力使高质量测评更易获取,惠及教师与学生。

原文摘要 · Abstract (English)

While large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instruction. One promising application is the generation of customized exams, tailored to specific course content. There has been significant recent excitement on automatically generating questions using artificial intelligence, but also comparatively little work evaluating the psychometric quality of these items in real-world educational settings. Filling this gap is an important step toward understanding generative AI's role in effective test design. In this study, we introduce and evaluate an iterative refinement strategy for question generation, repeatedly producing, assessing, and improving questions through cycles of LLM-generated critique and revision. We evaluate the quality of these AI-generated questions in a large-scale field study involving 91 classes -- covering computer science, mathematics, chemistry, and more -- in dozens of colleges across the United States, comprising nearly 1700 students. Our analysis, based on item response theory (IRT), suggests that for students in our sample the AI-generated questions performed comparably to expert-created questions designed for standardized exams. Our results illustrate the power of AI to make high-quality assessments more readily available, benefiting both teachers and students.

AI出题教育技术测试评估LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。