arXiv:2412.14670cs.CLcs.AI2024-12被引 1

剖析BERT如何理解动词+短语搭配的语法结构

Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT

  • 用MDS和GDV分析BERT各层对动词短语的表征效果
  • 中间层对语法结构的捕捉最准确,不同动词类别表现差异大
  • 适合研究语言模型内部机制的计算语言学工作者

本研究探究基于Transformer的大语言模型(如BERT)对动词-短语搭配(如'agree on'、'come back'、'give up')的内部表征能力,分析其在不同神经网络层中对词汇与句法细微差别的捕捉情况。研究基于英国国家语料库构建数据集,通过多维缩放(MDS)与广义区分值(GDV)计算进行模型训练与输出分析。结果表明,BERT的中间层最有效捕捉语法结构,且不同动词类别的表征准确率存在显著差异。这一发现挑战了神经网络处理语言元素时同质性的传统假设,揭示了网络架构与语言表征间的复杂互动关系。研究深化了对深度学习模型语言理解机制的认识,为优化神经架构以提升语言分析精度提供了新思路。

原文摘要 · Abstract (English)

This study investigates the internal representations of verb-particle combinations within transformer-based large language models (LLMs), specifically examining how these models capture lexical and syntactic nuances at different neural network layers. Employing the BERT architecture, we analyse the representational efficacy of its layers for various verb-particle constructions such as 'agree on', 'come back', and 'give up'. Our methodology includes a detailed dataset preparation from the British National Corpus, followed by extensive model training and output analysis through techniques like multi-dimensional scaling (MDS) and generalized discrimination value (GDV) calculations. Results show that BERT's middle layers most effectively capture syntactic structures, with significant variability in representational accuracy across different verb categories. These findings challenge the conventional uniformity assumed in neural network processing of linguistic elements and suggest a complex interplay between network architecture and linguistic representation. Our research contributes to a better understanding of how deep learning models comprehend and process language, offering insights into the potential and limitations of current neural approaches to linguistic analysis. This study not only advances our knowledge in computational linguistics but also prompts further research into optimizing neural architectures for enhanced linguistic precision.

语言模型语法结构表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。