通过合成数据检验多语言模型对主谓一致的语法捕捉能力
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement
- 构建多语言语法合成数据集,聚焦主谓一致现象
- 模型在相近语言间仍表现差异,语法结构未充分共享
- 适合研究多语言表示与语法归纳的学者参考
本文旨在探究多语言预训练语言模型在跨语言背景下捕捉抽象语言表征的程度。我们采用大规模、定制化的合成数据,结合新设计的多项选择任务和数据集Blackbird Language Matrices(BLMs),聚焦多种句式下的主谓一致这一具体语法结构现象。解决该任务需系统识别文本表征中的复杂语言模式与范式。通过两级架构——先在单句中检测句法对象及其属性,再在句子序列中发现规律——我们发现,尽管模型在一致的多语言文本上训练,但其在不同语言间仍存在显著差异,语法结构并未在语言间充分共享,甚至在关系密切的语言之间亦如此。
原文摘要 · Abstract (English)
In this paper, our goal is to investigate to what degree multilingual pretrained language models capture cross-linguistically valid abstract linguistic representations. We take the approach of developing curated synthetic data on a large scale, with specific properties, and using them to study sentence representations built using pretrained language models. We use a new multiple-choice task and datasets, Blackbird Language Matrices (BLMs), to focus on a specific grammatical structural phenomenon -- subject-verb agreement across a variety of sentence structures -- in several languages. Finding a solution to this task requires a system detecting complex linguistic patterns and paradigms in text representations. Using a two-level architecture that solves the problem in two steps -- detect syntactic objects and their properties in individual sentences, and find patterns across an input sequence of sentences -- we show that despite having been trained on multilingual texts in a consistent manner, multilingual pretrained language models have language-specific differences, and syntactic structure is not shared, even across closely related languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。