用多智能体框架生成认知多样化的选择题,提升题目质量与区分度。
Cognitively Diverse Multiple-Choice Question Generation: A Hybrid Multi-Agent Framework with Large Language Models
- 分步拆解出题任务,结合大模型与规则组件协同生成
- 生成题目更难、区分度更高,与阅读能力关联更强
- 适合教育测评、自动出题系统研发者使用
大语言模型(LLM)使自动化多项选择题(MCQ)生成成为可能,但可靠生成满足特定认知需求的题目仍具挑战。为此,我们提出ReQUESTA——一种混合式多智能体框架,可系统生成针对文本理解、推理和主旨把握的认知多样性题目。该框架将出题过程分解为专业化子任务,通过大模型驱动的智能体与基于规则的组件协同完成规划、可控生成、迭代评估与后处理。我们在学术说明文上开展大规模阅读理解研究,对比ReQUESTA生成的题目与单次提示的GPT-5零样本基线。心理测量分析显示,学习者答题结果表明ReQUESTA生成题目具有更高难度和更强区分度,且与整体阅读表现更紧密相关。专家评估进一步显示,题目在核心概念对齐性、干扰项语言一致性与语义合理性方面表现更优,尤其在推理类题目中优势显著。结果表明,混合式智能体编排能显著提升基于大模型生成的可靠性与可控性,凸显工作流设计在结构化内容生成中的关键作用。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have made automated multiple-choice question (MCQ) generation increasingly feasible; however, reliably producing items that satisfy controlled cognitive demands remains a challenge. To address this gap, we introduce ReQUESTA, a hybrid, multi-agent framework for generating cognitively diverse MCQs that systematically target text-based, inferential, and main idea comprehension. ReQUESTA decomposes MCQ authoring into specialized subtasks and coordinates LLM-powered agents with rule-based components to support planning, controlled generation, iterative evaluation, and post-processing. We evaluated the framework in a large-scale reading comprehension study using academic expository texts, comparing ReQUESTA-generated MCQs with those produced by a single-pass GPT-5 zero-shot baseline. Psychometric analyses of learner responses assessed item difficulty and discrimination, while expert raters evaluated question quality across multiple dimensions, including topic relevance and distractor quality. Results showed that ReQUESTA-generated items were consistently more challenging, more discriminative, and more strongly aligned with overall reading comprehension performance. Expert evaluations further indicated stronger alignment with central concepts and superior distractor linguistic consistency and semantic plausibility, particularly for inferential questions. These findings demonstrate that hybrid, agentic orchestration can systematically improve the reliability and controllability of LLM-based generation, highlighting workflow design as a key lever for structured artifact generation beyond single-pass prompting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。