用练习数据估算语言能力,精度媲美正式测试。
Implicit assessment of language learning during practice as accurate as explicit testing
- 将练习中的语言行为映射为项目,用IRT模型建模
- 实测显示练习推断能力与测试结果高度一致
- 适合想减少测试负担的智能辅导系统
智能辅导系统中对学习者水平的评估至关重要。本文在计算机辅助语言学习中应用项目反应理论(IRT),分别在测试和练习会话中评估学生能力。全面测试虽能提供精准画像,但存在诸多不便。因此,我们首先用不完美条件下的大规模测试数据训练IRT模型,用于指导高效且准确的自适应测试。仿真与真实数据实验均验证了该方法的高效性与准确性。其次,探索能否直接从练习过程中的数据估算能力,无需专门测试。通过将练习关联至语言构念,将其视为IRT中的‘题目’,实现建模。基于数千名学习者的大型研究结果显示,以教师评估为‘真实值’对比,仅基于练习数据的估计结果与测试结果相当,证实该方法可实现高精度能力推断。
原文摘要 · Abstract (English)
Assessment of proficiency of the learner is an essential part of Intelligent Tutoring Systems (ITS). We use Item Response Theory (IRT) in computer-aided language learning for assessment of student ability in two contexts: in test sessions, and in exercises during practice sessions. Exhaustive testing across a wide range of skills can provide a detailed picture of proficiency, but may be undesirable for a number of reasons. Therefore, we first aim to replace exhaustive tests with efficient but accurate adaptive tests. We use learner data collected from exhaustive tests under imperfect conditions, to train an IRT model to guide adaptive tests. Simulations and experiments with real learner data confirm that this approach is efficient and accurate. Second, we explore whether we can accurately estimate learner ability directly from the context of practice with exercises, without testing. We transform learner data collected from exercise sessions into a form that can be used for IRT modeling. This is done by linking the exercises to {\em linguistic constructs}; the constructs are then treated as "items" within IRT. We present results from large-scale studies with thousands of learners. Using teacher assessments of student ability as "ground truth," we compare the estimates obtained from tests vs. those from exercises. The experiments confirm that the IRT models can produce accurate ability estimation based on exercises.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。