用大模型给WordNet标注语言等级,助力外语学习者精准选词
CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning
- 用大模型匹配词义与欧洲语言框架等级,自动化标注词义难度
- 自建语料库训练的分类器达0.81宏平均F1,接近人工标注效果
- 适合语言教育研究者和NLP开发者,推动智能语言学习工具落地
尽管WordNet因其结构化的语义网络和丰富的词汇量而具有重要价值,但其细粒度的词义区分对第二语言学习者而言仍具挑战。为解决此问题,我们开发了基于欧洲共同语言参考框架(CEFR)的词义标注版WordNet,将语义网络与语言能力水平相结合。通过大语言模型计算WordNet中词义定义与英语词汇水平在线数据库条目间的语义相似度,实现标注自动化。为验证方法有效性,我们构建了一个大规模语料库,包含词义与对应CEFR等级信息,并基于该语料库开发上下文词汇分类器。实验表明,微调后的模型性能与使用金标准标注数据训练的模型相当。进一步结合金标准数据后,所构建分类器达到0.81的宏平均F1得分,间接证明转移标签与人工标注高度一致。所开发的标注版WordNet、语料库及分类器均已公开,旨在弥合自然语言处理与语言教育之间的鸿沟,促进更高效的语言学习。
原文摘要 · Abstract (English)
Although WordNet is a valuable resource because of its structured semantic networks and extensive vocabulary, its fine-grained sense distinctions can be challenging for second-language learners. To address this issue, we developed a version of WordNet annotated with the Common European Framework of Reference for Languages (CEFR), integrating its semantic networks with language-proficiency levels. We automated this process using a large language model to measure the semantic similarity between sense definitions in WordNet and entries in the English Vocabulary Profile Online. To validate our approach, we constructed a large-scale corpus containing both sense and CEFR-level information from the annotated WordNet and used it to develop contextual lexical classifiers. Our experiments demonstrate that models fine-tuned on this corpus perform comparably to those fine-tuned on gold-standard annotations. Furthermore, by combining this corpus with the gold-standard data, we developed a practical classifier that achieves a Macro-F1 score of 0.81. This result provides indirect evidence that the transferred labels are largely consistent with the gold-standard levels. The annotated WordNet, corpus, and classifiers are publicly available to help bridge the gap between natural language processing and language education, thereby facilitating more effective and efficient language learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。