研究BERT模型对习语和短语结构的注意力差异,揭示任务类型影响模型关注重点。
Attention on Multiword Expressions: A Multilingual Study of BERT-based Models with Regard to Idiomaticity and Microsyntax
- 对比不同任务微调后模型对习语与短语结构的注意力分布模式
- 语义任务微调使习语注意力在各层更均匀,语法任务微调提升底层对短语结构的关注
- 适用于自然语言处理中多语言语义与句法理解的研究者
本研究分析基于BERT架构的编码器模型在六种印欧语系语言(英语、德语、荷兰语、波兰语、俄语、乌克兰语)中,对两类多词表达(MWEs)——习语与微观句法单元(MSUs)——的注意力模式。习语具有语义非组合性,而MSUs表现出不符合标准语法分类的非常规句法行为。我们考察了在特定任务上微调后的模型如何分配对MWE的注意力,以及这种分配在语义任务与句法任务间的差异。结果表明,微调显著影响模型对多词表达的注意力机制:语义任务微调的模型对习语的注意力在各层分布更均匀;而句法任务微调的模型在低层增加了对MSUs的关注,符合句法处理需求。
原文摘要 · Abstract (English)
This study analyzes the attention patterns of fine-tuned encoder-only models based on the BERT architecture (BERT-based models) towards two distinct types of Multiword Expressions (MWEs): idioms and microsyntactic units (MSUs). Idioms present challenges in semantic non-compositionality, whereas MSUs demonstrate unconventional syntactic behavior that does not conform to standard grammatical categorizations. We aim to understand whether fine-tuning BERT-based models on specific tasks influences their attention to MWEs, and how this attention differs between semantic and syntactic tasks. We examine attention scores to MWEs in both pre-trained and fine-tuned BERT-based models. We utilize monolingual models and datasets in six Indo-European languages - English, German, Dutch, Polish, Russian, and Ukrainian. Our results show that fine-tuning significantly influences how models allocate attention to MWEs. Specifically, models fine-tuned on semantic tasks tend to distribute attention to idiomatic expressions more evenly across layers. Models fine-tuned on syntactic tasks show an increase in attention to MSUs in the lower layers, corresponding with syntactic processing requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。