系统分析337篇论文,揭示大模型的句法知识现状与局限。
The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models
- 综合3000+数据点,评估不同语言模型在句法任务中的表现
- 模型对形式化句法掌握良好,但跨句法语义界面表现较弱
- 研究集中于英语和BERT类模型,方法不统一限制机制理解
我们对337篇评估基于Transformer的语言模型(TLMs)句法能力的论文进行了系统综述,涵盖超过3000个数据点,覆盖多种句法现象、语言、模型和方法。数据整体表明,TLMs编码了非平凡量的句法知识。行为证据显示,模型在形式化句法现象上表现强劲,但在句法-语义接口现象上表现较弱且波动较大。语言数字化支持不足时,性能也持续偏低。探针与机制研究进一步验证了句法知识的存在。然而,由于多数研究仍为观察性且方法异质性强,对句法处理背后计算机制的理解依然有限。同时,文献高度集中于英语和BERT类模型。本文讨论结果影响,并提出未来研究建议。
原文摘要 · Abstract (English)
We present a systematic review of 337 articles evaluating the syntactic abilities of Transformer-based language models (TLMs), reporting on over 3,000 datapoints spanning a wide range of syntactic phenomena, languages, models, and methods. We take the data to collectively show that TLMs encode a non-trivial amount of syntactic knowledge. Behavioral evidence shows strong performance on formal syntactic phenomena, but weaker and more variable performance on phenomena at the syntax-semantics interface. Performance is also consistently lower for languages with less digital support. Probing and mechanistic studies further support the presence of syntactic knowledge in TLMs. Yet, because most work remains observational and methodologically heterogeneous, insight into the detailed computational mechanisms underlying syntactic processing remains limited. At the same time, the literature remains heavily concentrated on English and BERT-like models. We discuss the implications of our results and provide recommendations for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。