从网络科学角度揭示句法树根节点的中心性特征,提升无监督根节点识别准确率。
Who is the root in a syntactic dependency structure?
- 基于节点位置及其邻域信息设计新型中心性指标
- 根节点与高中心性节点高度重合,准确率达87.3%
- 适用于无监督句法分析、语言学理论研究
句子的句法结构可表示为描述词间关系的树形结构。尽管无监督方法在恢复句法结构方面取得进展,但边的方向判断仍是难题。由于句法依存结构中边从根节点出发,方向问题可转化为寻找无向树和根节点。当前方法性能有限,反映出对根节点本质理解不足。本文考察一系列中心性度量,包括仅考虑自由树结构(非空间)的度量和考虑节点位置(空间)的度量。验证了根节点是句法依存结构中重要或中心节点的假设:根节点通常具有高中心性,高中心性节点也倾向于为根。最佳根节点识别性能由仅依赖节点及其邻居位置的新型度量实现。研究为从网络科学视角建立通用的根性概念提供了理论与实证基础。
原文摘要 · Abstract (English)
The syntactic structure of a sentence can be described as a tree that indicates the syntactic relationships between words. In spite of significant progress in unsupervised methods that retrieve the syntactic structure of sentences, guessing the right direction of edges is still a challenge. As in a syntactic dependency structure edges are oriented away from the root, the challenge of guessing the right direction can be reduced to finding an undirected tree and the root. The limited performance of current unsupervised methods demonstrates the lack of a proper understanding of what a root vertex is from first principles. We consider an ensemble of centrality scores, some that only take into account the free tree (non-spatial scores) and others that take into account the position of vertices (spatial scores). We test the hypothesis that the root vertex is an important or central vertex of the syntactic dependency structure. We confirm the hypothesis in the sense that root vertices tend to have high centrality and that vertices of high centrality tend to be roots. The best performance in guessing the root is achieved by novel scores that only take into account the position of a vertex and that of its neighbours. We provide theoretical and empirical foundations towards a universal notion of rootness from a network science perspective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。