将词向量转为可解释的语法表示,构建层次化词向量
Interpretable Syntactic Representations Enable Hierarchical Word Vectors
- 用语法结构压缩密集向量,生成紧凑可解释表示
- 层次化词向量在基准测试中优于原始向量
- 适合需要可解释性的自然语言处理研究者
当前的分布式词向量是密集且不可解释的,导致其语义解释相对、冗余且难以理解。本文提出一种方法,将词向量转换为简约的语法表示,使结果更紧凑、可解释,便于可视化与对比,且其解释符合人类判断。基于预训练向量,采用类似人类学习的层级机制,逐步构建层次化词向量。该生成过程与学习方式计算高效。最重要的是,语法表示提供了合理的向量解释,后续的层次化向量在基准测试中表现优于原始向量。
原文摘要 · Abstract (English)
The distributed representations currently used are dense and uninterpretable, leading to interpretations that themselves are relative, overcomplete, and hard to interpret. We propose a method that transforms these word vectors into reduced syntactic representations. The resulting representations are compact and interpretable allowing better visualization and comparison of the word vectors and we successively demonstrate that the drawn interpretations are in line with human judgment. The syntactic representations are then used to create hierarchical word vectors using an incremental learning approach similar to the hierarchical aspect of human learning. As these representations are drawn from pre-trained vectors, the generation process and learning approach are computationally efficient. Most importantly, we find out that syntactic representations provide a plausible interpretation of the vectors and subsequent hierarchical vectors outperform the original vectors in benchmark tests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。