arXiv:2607.07141cs.CLstat.ME2026-07

用文本嵌入预测试题参数,揭示了可预测性上限与评估盲区。

From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings

论文配图:From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings
图 1 · 摘自论文原文
  • 基于文本嵌入的正则化回归预测试题难度,效果显著。
  • 难度参数可解释57%~63%的可靠方差,但猜测参数几乎不可预测。
  • 提出可靠性与设计上限,为测评基准提供新评估视角。

新题目的心理测量属性通常需实地测试才能确定,导致题目校准存在冷启动问题。从特征预测题目参数是长期存在的测量难题,可追溯至线性逻辑测验模型;现代文本嵌入可自动构建传统手工指定的设计矩阵。本文提出一个评估框架,包含对题目文本嵌入的正则化回归、重复交叉验证的决定系数(报告其重采样标准差),以及两个性能上限:基于参数标准误的可靠性上限,和基于模拟功效校准的设计上限。在数学题库(EEDI)和医学执照基准(BEA 2024)上的应用表明,题目难度可高度由文本预测(重复交叉验证R² = 0.53,约达到其可靠性上限的57%),而区分度和伪猜测参数则较难预测。然而,将结果与上限对比后发现,这一明显差异源于目标可靠性而非文本信号强度:文本对难度目标统一能恢复57%至63%的可靠方差,而3PL模型的伪猜测参数可靠性上限接近零,使其在当前精度下不可行。在BEA上,基于嵌入的回归虽解释方差极少,却匹配排行榜RMSE,凸显量纲无关指标与明确上限在基准评估中的关键作用。最后,我们证明单次训练测试分割会使R²虚高0.1至0.15,强调重复交叉验证对校准支持及未来基准建设的必要性。

原文摘要 · Abstract (English)

Newly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold start problem for item calibration. Predicting item parameters from features is a long standing measurement problem dating back to the Linear Logistic Test Model; modern text embeddings now automate the design matrices traditionally specified by hand. We propose an evaluation framework combining regularized regression on item text embeddings, repeated cross validated R squared reported with its resampling standard deviation, and two performance upper bounds: a reliability ceiling derived from parameter standard errors, and a design ceiling derived from simulation based power calibration. Applying this framework to a mathematics item bank (EEDI) and a medical licensure benchmark (BEA 2024), we find that item difficulty is highly predictable from text (repeated cross validated R squared = 0.53, or about 57% of its reliability ceiling), whereas discrimination and pseudo guessing appear less predictable. However, evaluating these results against our ceilings reveals that this apparent hierarchy stems from target reliability rather than text signal strength: text uniformly recovers 57 to 63% of the reliable variance across difficulty targets, whereas the 3PL pseudo guessing parameter has a reliability ceiling near zero, making it an unviable target at current precision. On BEA, embedding based regression matches leaderboard RMSE despite explaining almost no variance, highlighting the critical need for scale free metrics and explicit ceilings in benchmarking. Finally, we show that a single train and test split can inflate apparent accuracy by 0.1 to 0.15 in R squared, underscoring the necessity of repeated cross validation for calibration support applications and future benchmark construction.

试题生成文本嵌入心理测量评估上限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。