通过结构化建模选择题选项,显著提升自动难度预测准确率。
Structure-Aware Modeling of Multiple-Choice Questions Improves Automatic Difficulty Estimation
- 将题干和干扰项分开展示,用位置标签或求和方式融合信息。
- 在自然科学题上达到R²=0.83,社会科学题R²=0.71。
- 不依赖顺序的模型参数减半,适合大规模教育应用。
自动题目难度估计(AQDE)在教育评估中潜力巨大,可提供媲美专家判断的难度预测,同时降低试测的时间与成本,并支持数字化测试扩展。以往研究对将干扰项作为额外文本加入题干和正确答案是否持续提升预测效果存在分歧。我们假设干扰项信息的有效性取决于其结构表示方式,显式建模干扰项为独立组件可优于忽略该信息的基线模型。为此,我们设计了受控架构,将多选题各组成部分作为独立输入,以分离干扰项内容与顺序的影响。具体地,每个干扰项被编码为独立文本输入,其表示通过带位置标记的拼接或无序求和进行聚合。我们在两个智利数据集(2016–2020年自然科学与社会科学,共4,114道题)上评估,结果表明,最佳的干扰项感知架构在自然科学题上达到R²=0.83,社会科学题上R²=0.71,显著优于仅使用题干和正确答案的基线模型。一种无序不变的变体在参数量减少约一半的同时保持相近精度,展现出良好的准确性-效率权衡。结果表明,结构化信息(尤其是干扰项内容)是提升预测性能的关键,支持开发高效、结构感知的模型用于大规模教育场景。
原文摘要 · Abstract (English)
Automatic Question Difficulty Estimation (AQDE) holds growing promise for educational assessment because it has the potential to yield difficulty estimates that are competitive with expert judgment, while helping reduce the time and financial burden associated with pilot administrations and scaling to digital testing contexts. Prior AQDE studies report mixed evidence on whether adding distractors as additional text to the question stem and the correct key consistently improves difficulty prediction. We hypothesize that the effectiveness of distractor information depends on its structural representation, and that explicitly modeling distractors as separate components improves difficulty estimation over baselines that omit this information. To address this, we designed controlled architectures that model MCQ components as distinct inputs to isolate the contribution of distractor content and order. Specifically, we represented distractors by encoding each distractor as its own text input and aggregating their representations either with order-aware concatenation (with positional tags) or with an order-invariant summation. We evaluated these architectures using two Chilean datasets (Natural and Social Sciences, 2016-2020; 4,114 multiple-choice questions). Compared to a simpler model that only used the question stem and the key, our best distractor-aware architecture achieved higher predictive performance, reaching R^2 = 0.83 for Natural Sciences and R^2 = 0.71 for Social Sciences items. An order-invariant variant achieved nearly the same accuracy with approximately half as many parameters, offering a favorable accuracy-efficiency trade-off. These results show that structural information (especially distractor content) drives gains in predictive accuracy, supporting the development of efficient, structure-aware models that are computationally viable for large-scale educational applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。