arXiv:1301.37810cs.CL2013-01

用更少计算量快速训练出高质量词向量,效果超越以往方法。

Efficient Estimation of Word Representations in Vector Space

  • 提出两种新模型架构,从大规模语料中高效生成词向量
  • 16亿词数据下一天内完成训练,准确率显著提升
  • 在句法与语义相似度任务上达到当时最优表现

我们提出了两种新颖的模型架构,用于从超大规模语料中计算连续的词向量表示。通过词相似度任务评估表示质量,并与此前基于不同类型神经网络的最佳方法进行比较。结果显示,在远低于以往成本的计算开销下,准确率有显著提升:仅需不到一天时间即可从16亿词的数据集中学到高质量的词向量。此外,这些向量在作者测试集上的句法与语义相似度任务中表现出当前最优性能。

原文摘要 · Abstract (English)

We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.

词向量自然语言处理深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。