首个吉尔吉斯语词嵌入评估数据集,助力自然语言处理模型选型。
HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings
- 构建首个吉尔吉斯语词向量评估的‘银标准’数据集
- 通过专家标注相似性验证向量距离有效性
- 适合低资源语言研究者与词嵌入质量评估人员
现代应用计算语言学的关键任务之一是构建词向量表示(词嵌入),广泛应用于情感分析、信息抽取等自然语言处理任务。为选择合适的词嵌入生成方法,常需进行质量评估。标准方法是计算具有专家评估‘相似性’的词向量间距离。本文首次提出吉尔吉斯语的‘银标准’评估数据集,同时训练相应模型,并通过质量评估指标验证数据集适用性。
原文摘要 · Abstract (English)
One of the key tasks in modern applied computational linguistics is constructing word vector representations (word embeddings), which are widely used to address natural language processing tasks such as sentiment analysis, information extraction, and more. To choose an appropriate method for generating these word embeddings, quality assessment techniques are often necessary. A standard approach involves calculating distances between vectors for words with expert-assessed 'similarity'. This work introduces the first 'silver standard' dataset for such tasks in the Kyrgyz language, alongside training corresponding models and validating the dataset's suitability through quality evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。