用深度学习自动挑选适合高中生词汇教学的优质语境。
Predicting Contextual Informativeness for Vocabulary Learning using Deep Learning
- 用微调后的Qwen3嵌入+非线性回归预测语境价值。
- 最优模型仅丢弃70%优质语境,好坏比达440:1。
- 适合教育科技、智能教学系统开发者参考。
我们提出一种现代深度学习系统,可自动识别适用于高中生母语词汇教学的优质上下文实例。对比三种建模方法:(i) 基于MPNet统一上下文嵌入的无监督相似度策略;(ii) 使用指令感知的微调Qwen3嵌入并搭配非线性回归头的有监督框架;(iii) 在方法(ii)基础上加入手工设计的语境特征。引入新的评估指标——保留能力曲线,可视化丢弃优质语境比例与好坏语境比之间的权衡,提供统一性能视角。方法(iii)表现最佳,实现440:1的好坏语境比,同时仅丢弃70%的优质语境。结果表明,经人类监督引导的现代神经网络嵌入模型,可低成本生成大量近乎完美的教学语境,适用于多种目标词汇。
原文摘要 · Abstract (English)
We describe a modern deep learning system that automatically identifies informative contextual examples (\qu{contexts}) for first language vocabulary instruction for high school student. Our paper compares three modeling approaches: (i) an unsupervised similarity-based strategy using MPNet's uniformly contextualized embeddings, (ii) a supervised framework built on instruction-aware, fine-tuned Qwen3 embeddings with a nonlinear regression head and (iii) model (ii) plus handcrafted context features. We introduce a novel metric called the Retention Competency Curve to visualize trade-offs between the discarded proportion of good contexts and the \qu{good-to-bad} contexts ratio providing a compact, unified lens on model performance. Model (iii) delivers the most dramatic gains with performance of a good-to-bad ratio of 440 all while only throwing out 70\% of the good contexts. In summary, we demonstrate that a modern embedding model on neural network architecture, when guided by human supervision, results in a low-cost large supply of near-perfect contexts for teaching vocabulary for a variety of target words.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。