arXiv:2605.18562stat.MEcs.AI2026-05被引 1

用大模型估算新题难度,省去试测成本。

Estimating Item Difficulty with Large Language Models as Experts

论文配图:Estimating Item Difficulty with Large Language Models as Experts
图 1 · 摘自论文原文
  • 让大模型通过对比或直接打分判断题目难易
  • 在数学题上相关性达中等到强,部分接近人工水平
  • 配对比较+概率提示效果最佳,适合教育产品快速出题

准确估计题目难度对有效评估和自适应学习至关重要。然而,新题缺乏答题数据,传统预测试和专家评分成本高、耗时长,而机器学习方法通常需要大量标注数据。近期研究显示大语言模型(LLMs)可能有帮助,但其模拟专家的评估流程与提示设计尚不明确。本研究评估了三种现成的LLM作为无响应数据条件下新题难度评估者的表现。基于在线学习系统中的题库,研究涵盖小学数学6个领域,以实证难度为参考标准。采用全因子实验设计,考察三种因素:评分格式(绝对值 vs 配对比较)、决策类型(硬判断 vs 基于标记概率的估计)和提示策略(零样本 vs 少样本)。使用斯皮尔曼等级相关系数比较模型估计与实证难度。结果显示,跨领域中LLM估计与实证难度呈中度至强正相关;对于简单算术题,部分配置达到以往人类专家研究报道的准确率上限。配对比较在未额外优化时表现优于绝对判断;但当引入标记级概率并提供已知难度示例时,绝对判断也实现中到高度一致性。研究表明LLM是初始题目校准的有力工具,并提供了高效工作流配置建议。

原文摘要 · Abstract (English)

Accurate estimates of item difficulty are essential for valid assessment and effective adaptive learning. However, for newly created tasks, response data are typically unavailable. Pretesting and expert judgement can be costly and slow, while machine learning methods often require large labelled training datasets. Recent work suggests that large language models (LLMs) may help. However, there is limited evidence on the elicitation procedures and prompt configurations used to emulate experts for difficulty estimation. This study addresses this gap by evaluating three off-the-shelf LLMs as difficulty raters for newly created items without access to response data. Using an item bank from an online learning system, the study examined 6 domains of primary-school mathematics, with empirical difficulty estimates treated as empirical reference. The study used a full factorial design crossing three factors: judgement format (absolute vs pairwise), decision type (hard decisions vs token-probability-based estimates), and prompting strategy (zero-shot vs few-shot). LLM-derived difficulty estimates were compared with empirical difficulties using Spearman rank correlations. Across domains, LLM-based estimates exhibited moderate to strong positive correlations with empirical item difficulties. For simpler arithmetic tasks, some configurations approached the upper end of the accuracy range reported for human experts in previous research. Pairwise comparison consistently outperformed absolute judgement in the absence of additional refinements. However, when token-level probabilities were incorporated and examples of items with known empirical difficulty were provided, the absolute judgement configuration likewise demonstrated moderate-to-high alignment. The study positions LLMs as a promising tool for initial item calibration and offers insights into effective workflow configuration.

大模型题目难度教育AILLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。