提出新模型与评估方法,更准确地判断题目难度等级。
Ordinality in Discrete-level Question Difficulty Estimation: Introducing Balanced DRPS and OrderedLogitNN
- 引入神经网络版有序逻辑回归模型OrderedLogitNN
- 在RACE++和ARC数据集上表现优于传统方法
- 设计平衡型评分指标,解决类别不平衡问题
近年来,自然语言处理技术被广泛用于题目难度估计(QDE)。题目难度通常以离散等级表示,具有从易到难的有序结构,应归为有序回归任务。然而,现有研究多忽略这一特性,采用分类或离散回归模型,未充分探索专用有序回归方法。同时,评估指标与建模范式绑定,缺乏可比性;部分指标忽视难度等级的有序性,且均未有效处理类别不平衡,导致性能评估偏差。本研究通过基准测试三种模型输出——离散回归、分类与有序回归,引入新的平衡离散排名概率分数(Balanced DRPS),该指标同时捕捉有序性和类别不平衡。除使用经典有序回归方法外,还提出将经济学中的有序逻辑回归拓展至神经网络的OrderedLogitNN。在RACE++和ARC数据集上微调BERT后发现,OrderedLogitNN在复杂任务中表现显著更优。平衡DRPS为离散级题目难度估计提供了稳健公平的评估标准,为后续研究奠定基础。
原文摘要 · Abstract (English)
Recent years have seen growing interest in Question Difficulty Estimation (QDE) using natural language processing techniques. Question difficulty is often represented using discrete levels, framing the task as ordinal regression due to the inherent ordering from easiest to hardest. However, the literature has neglected the ordinal nature of the task, relying on classification or discretized regression models, with specialized ordinal regression methods remaining unexplored. Furthermore, evaluation metrics are tightly coupled to the modeling paradigm, hindering cross-study comparability. While some metrics fail to account for the ordinal structure of difficulty levels, none adequately address class imbalance, resulting in biased performance assessments. This study addresses these limitations by benchmarking three types of model outputs -- discretized regression, classification, and ordinal regression -- using the balanced Discrete Ranked Probability Score (DRPS), a novel metric that jointly captures ordinality and class imbalance. In addition to using popular ordinal regression methods, we propose OrderedLogitNN, extending the ordered logit model from econometrics to neural networks. We fine-tune BERT on the RACE++ and ARC datasets and find that OrderedLogitNN performs considerably better on complex tasks. The balanced DRPS offers a robust and fair evaluation metric for discrete-level QDE, providing a principled foundation for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。