arXiv:2411.16993cs.CL2024-11被引 3

Tree Transformer虽有结构化设计,但实际未能有效学习句法成分结构。

Tree Transformers are an Ineffective Model of Syntactic Constituency

  • 用改进注意力机制构建树状结构以引导句法成分
  • 预训练后树结构缺乏有意义的语法表征,任务表现仅微弱提升
  • 对语言学中的递归成分结构建模效果有限,不适合依赖句法结构的任务

语言学家长期认为自然语言语法的核心特征是语言单位的递归组成结构。现有研究表明,当前最先进的语言模型缺乏对这一特征的内在偏置。为此,提出了包括树形变压器(Tree Transformer)在内的多种替代模型,通过修改注意力机制来引入成分结构的归纳偏置。本文通过对大型树形变压器进行语言建模预训练,研究其学习到的句子成分树表示,发现缺乏有意义的结构证据。进一步在需要成分结构的错误检测任务中评估其与同类Transformer模型的表现,结果表明树形变压器虽略有优势,但无明显实质性提升。总体而言,目前缺乏支持树形变压器作为句法成分有效模型的充分证据。

原文摘要 · Abstract (English)

Linguists have long held that a key aspect of natural language syntax is the recursive organization of language units into constituent structures, and research has suggested that current state-of-the-art language models lack an inherent bias towards this feature. A number of alternative models have been proposed to provide inductive biases towards constituency, including the Tree Transformer, which utilizes a modified attention mechanism to organize tokens into constituents. We investigate Tree Transformers to study whether they utilize meaningful and/or useful constituent structures. We pretrain a large Tree Transformer on language modeling in order to investigate the learned constituent tree representations of sentences, finding little evidence for meaningful structures. Next, we evaluate Tree Transformers with similar transformer models on error detection tasks requiring constituent structure. We find that while the Tree Transformer models may slightly outperform at these tasks, there is little evidence to suggest a meaningful improvement. In general, we conclude that there is little evidence to support Tree Transformer as an effective model of syntactic constituency.

句法结构树形模型语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。