更新了2024版英文词向量模型,更贴合当前语言使用。
A New Pair of GloVes
- 用Wikipedia、Gigaword和Dolma子集训练新词向量
- 在新词汇和非西方新闻实体识别上表现更好
- 修复了原版模型数据来源不清晰的问题
本文介绍了2024年新版英语GloVe词向量模型。尽管2014年原版模型被广泛使用,但语言与世界持续演变,现有使用场景亟需更新。此外,原版模型的数据版本与预处理流程缺乏明确记录,本次工作加以补全。我们基于Wikipedia、Gigaword及Dolma的一个子集训练了两组词嵌入。通过词汇对比、直接测试和命名实体识别(NER)任务评估发现,2024版向量包含更多文化与语言相关的新兴词汇,在类比与相似性等结构任务上表现相当,并在近期时序敏感的NER数据集(如非西方新闻语料)中取得显著提升。
原文摘要 · Abstract (English)
This report documents, describes, and evaluates new 2024 English GloVe (Global Vectors for Word Representation) models. While the original GloVe models built in 2014 have been widely used and found useful, languages and the world continue to evolve and we thought that current usage could benefit from updated models. Moreover, the 2014 models were not carefully documented as to the exact data versions and preprocessing that were used, and we rectify this by documenting these new models. We trained two sets of word embeddings using Wikipedia, Gigaword, and a subset of Dolma. Evaluation through vocabulary comparison, direct testing, and NER tasks shows that the 2024 vectors incorporate new culturally and linguistically relevant words, perform comparably on structural tasks like analogy and similarity, and demonstrate improved performance on recent, temporally dependent NER datasets such as non-Western newswire data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。