用大模型+词汇表精细评估二语单词使用水平
Exploiting the English Vocabulary Profile for L2 word-level vocabulary assessment with LLMs
- 结合大模型与英语词汇表,分析句子中词汇的语境使用
- 在多义词和搭配词上表现优于传统词性基线
- 适合语言测评研究者和教育技术开发者
词汇使用是第二语言能力的核心。现有自动化评估多关注词的上下文无关或词性相关用法。本文提出新方法,利用大语言模型(LLMs)与英语词汇表(EVP)实现细粒度词汇评估。EVP可将词汇在具体语境中的使用关联到相应语言水平。研究评估了LLMs对学习者作文中单个词汇的水平判断能力,解决多义性、语境变化和固定搭配等挑战。对比词性基线,LLMs能利用更多语义信息,表现更优。还探讨了词汇级与作文级水平的相关性。最后用于检验EVP水平标注的一致性。结果表明,LLMs非常适合该词汇评估任务。
原文摘要 · Abstract (English)
Vocabulary use is a fundamental aspect of second language (L2) proficiency. To date, its assessment by automated systems has typically examined the context-independent, or part-of-speech (PoS) related use of words. This paper introduces a novel approach to enable fine-grained vocabulary evaluation exploiting the precise use of words within a sentence. The scheme combines large language models (LLMs) with the English Vocabulary Profile (EVP). The EVP is a standard lexical resource that enables in-context vocabulary use to be linked with proficiency level. We evaluate the ability of LLMs to assign proficiency levels to individual words as they appear in L2 learner writing, addressing key challenges such as polysemy, contextual variation, and multi-word expressions. We compare LLMs to a PoS-based baseline. LLMs appear to exploit additional semantic information that yields improved performance. We also explore correlations between word-level proficiency and essay-level proficiency. Finally, the approach is applied to examine the consistency of the EVP proficiency levels. Results show that LLMs are well-suited for the task of vocabulary assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。