揭示大模型语法探针的局限性:距离近就关联,深层结构难捕捉。
Probing Syntax in Large Language Models: Successes and Remaining Challenges
- 通过控制语料测试探针,发现词距近则更易被关联
- 对深层语法结构表示差,受名词交互和错误动词干扰
- 词义可预测性不影响探针表现,适合评估模型内部表征
大型语言模型(LLMs)的激活值中可轻松读取句子的句法结构。然而,用于揭示这一现象的句法探针通常在无筛选的句子集合上进行评估,因此尚不清楚句法或统计因素是否系统性地影响这些句法表征。为解决此问题,我们在三个受控基准上对句法探针进行了深入分析。结果表明:第一,句法探针存在表面偏差——句子中两个词越接近,越可能被判定为句法相关;第二,句法探针难以表示深层句法结构,且易受相互作用的名词或不合语法的动词形式干扰;第三,单个词的可预测性并未显著影响探针表现。本研究揭示了当前句法探针面临的挑战,并提供了由受控刺激构成的基准,以更准确评估其性能。
原文摘要 · Abstract (English)
The syntactic structures of sentences can be readily read-out from the activations of large language models (LLMs). However, the ``structural probes'' that have been developed to reveal this phenomenon are typically evaluated on an indiscriminate set of sentences. Consequently, it remains unclear whether structural and/or statistical factors systematically affect these syntactic representations. To address this issue, we conduct an in-depth analysis of structural probes on three controlled benchmarks. Our results are three-fold. First, structural probes are biased by a superficial property: the closer two words are in a sentence, the more likely structural probes will consider them as syntactically linked. Second, structural probes are challenged by linguistic properties: they poorly represent deep syntactic structures, and get interfered by interacting nouns or ungrammatical verb forms. Third, structural probes do not appear to be affected by the predictability of individual words. Overall, this work sheds light on the current challenges faced by structural probes. Providing a benchmark made of controlled stimuli to better evaluate their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。