arXiv:2607.28634cs.CLcs.LG2026-07

LLM预测试题难度能力有限,生成特定难度题目需谨慎。

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

  • 用提示工程测试多种LLM,发现GPT-4.1在零样本下表现最佳
  • 最高准确率QWK为0.578,仍低于ConvBERT的0.625
  • 模型普遍低估难题,且语义嵌入难以区分难度层级

试题难度估计在形成性评估和大规模高风险总结性评估中至关重要。本研究探讨了大语言模型(LLMs)在预测大型阅读写作测试题难度等级方面的表现,考察了多种提示策略与参数设置,并对比了编码器类语言模型及基于特征的监督机器学习模型。零样本条件下,温度设为0的GPT-4.1取得最高准确率,二次加权肯德尔协调系数(QWK)达0.578。然而,该表现仍低于ConvBERT模型(QWK=0.625),后者优于最优的特征型监督学习模型。进一步分析显示,所有LLM均难以识别难题,尤其是当前先进的GPT-5.4倾向于低估难度。嵌入空间降维表明,不同难度等级的试题嵌入高度混杂,说明仅靠试题语义信息不足以准确判断难度。研究结果表明,若LLM无法真正理解试题难度,且随着能力提升反而更倾向于将所有题目视为简单,则在生成目标难度试题时应保持审慎。

原文摘要 · Abstract (English)

The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study investigated various prompting strategies and parameter settings across multiple LLMs. LLM performance was compared with encoder-only language models and feature-based supervised machine learning models. Zero-shot GPT-4.1 with a temperature of 0 yielded the highest item difficulty level prediction accuracy, with a quadratic weighted kappa (QWK) of 0.578. However, LLMs' prediction accuracy was lower than that of ConvBERT (QWK = 0.625), which outperformed the best feature-based supervised machine learning model. Further analysis showed that all LLMs struggled to label hard items; in particular, the current advanced GPT-5.4 tended to underestimate item difficulty levels. Dimension reduction of embeddings showed that item embeddings from different difficulty levels were mixed together, indicating that semantic information from items alone is likely insufficient for item difficulty level prediction. The findings suggest that if LLMs cannot understand item difficulty levels as evidenced by empirical data and tend to treat most items as easy when their own capabilities increase, caution should be exercised when using LLMs to generate items with targeted difficulty levels.

试题生成难度预测LLM评估教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。