arXiv:2608.27813cs.CL2026-08

发现语言模型语法表征受词距与句法头多样性影响

Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy

论文配图:Representation of syntax in LLMs through the lens of linear distance and similarity-aware entropy
图 1 · 摘自论文原文
  • 用标签级无向依存得分拆解语法重建精度
  • 词距和句法头多样性可解释90%以上得分差异
  • 适用于不同规模架构的模型,揭示表征抽象性

结构探针由Hewitt和Manning提出,用于从神经语言模型的潜在表示中重构句法树。其评估方式为在标注语料上计算正确重建的句法树边比例(以无向无标签依存得分衡量)。本文将该指标细化,考察每类句法关系的无向依存得分(UASL),从而揭示不同句法关系间的差异,这些差异对应语言学上的区分。此外,我们识别出两个能解释大部分UASL变异性的因素:(i) 相关词之间的线性距离(对数尺度下的均值与离散度),(ii) 句法关系头的多样性(相似性感知熵)。这些结果在多种模型规模与架构下均成立,揭示了语言模型中语法表征的抽象程度及其对嵌入空间几何特性的依赖。

原文摘要 · Abstract (English)

Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are evaluated by calculating the proportion of syntactic tree edges correctly reconstructed over an annotated corpus (as measured by undirected unlabeled attachment score). Here, we disaggregate this measure, considering undirected attachment score by label (UASL), which assesses the reconstruction accuracy of each syntactic relation separately, establishing important differences among relations that overlap linguistic distinctions. Moreover, we identify two factors that predict most of UASL's variability across relations: (i) the mean and dispersion of the linear distance (on a log scale) between the related words, and (ii) the diversity (similarity-aware entropy) of the syntactic relation's head. These results, which hold across a range of model sizes and architectures, shed light on the degree of abstraction of the representation of syntax in language models and the dependence of such representation on geometric properties of the embedding space.

句法分析语言模型表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。