arXiv:2506.16678cs.CL2025-06EMNLP被引 9

探针发现的语法表征,无法预测模型在语法任务中的表现。

Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations

  • 用探针分析32个开源模型的语法表征机制
  • 探针准确率与下游语法任务表现无显著相关性
  • 适合关注模型可解释性与评估方法的读者

大型语言模型在文本处理与生成中展现出强大的语法掌握能力,暗示其内部可能已内化层级语法与依存关系。然而,这些模型如何表征语法结构仍是可解释性研究中的开放问题。探针方法试图识别激活值中线性编码的语法特征,但目前尚无全面研究验证探针准确率是否能可靠预测模型在下游语法任务中的表现。本研究采用“机制与结果”框架,评估了32个开源Transformer模型,发现通过探针提取的语法特征无法预测英语语言现象下的目标语法评估结果。结果表明,探针所揭示的潜在语法表示与下游任务中的可观测语法行为之间存在显著脱节。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit a robust mastery of syntax when processing and generating text. While this suggests internalized understanding of hierarchical syntax and dependency relations, the precise mechanism by which they represent syntactic structure is an open area within interpretability research. Probing provides one way to identify the mechanism of syntax being linearly encoded in activations, however, no comprehensive study has yet established whether a model's probing accuracy reliably predicts its downstream syntactic performance. Adopting a "mechanisms vs. outcomes" framework, we evaluate 32 open-weight transformer models and find that syntactic features extracted via probing fail to predict outcomes of targeted syntax evaluations across English linguistic phenomena. Our results highlight a substantial disconnect between latent syntactic representations found via probing and observable syntactic behaviors in downstream tasks.

语言模型可解释性探针语法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。