用词向量揭示词性分类的模糊边界,构建三维语义空间可视化词性关系。
Semantic Space of Parts of Speech

- 通过降维将词向量映射到三维空间,捕捉词性间的连续语义关系。
- 发现部分词处于词性边界,原型词与过渡词分布清晰可辨。
- 适用于语言学研究者及自然语言处理中词性消歧的改进方向。
词性标注在欧洲语言学传统中被视为明确分类,语料库语言学中每个词仅属一类词性。然而,这种分类主要依赖注释手册中的主观决策。由于某些词在语义或句法上介于不同词性之间,且部分词性本身更接近,词性分类本质上具有模糊性。本文利用word2vec词向量,训练神经网络将高维嵌入降至与词性相关的三维空间,将数千个词语映射至该空间,揭示哪些词为典型代表,哪些位于边界,并可视化词性间的关系。研究采用Universal Dependencies的法语、捷克语、芬兰语、俄语和英语词性标签。
原文摘要 · Abstract (English)
Parts of speech categorization is understood in the European linguistic tradition as crisp categorization, which is also reflected in corpus linguistics, where each disambiguated token is assigned exactly one POS. However, the assigned categories are largely determined by arbitrary decisions distilled into annotation manuals. Since some words stand between parts of speech in their semantics or typical syntax, and some parts of speech are closer to each other than others, POS categorization seems inherently fuzzy. We analyze this fuzziness using word2vec embeddings, training a neural network to reduce their high dimensionality to three dimensions relevant for determining parts of speech. This creates a three-dimensional space onto which we map several thousand words, revealing which are prototypical and which lie on the boundaries, and visualizing relationships between parts of speech. The study uses Universal Dependencies POS tags for French, Czech, Finnish, Russian, and English.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。