评估多语言模型在三种低资源语言中的句法能力,发现表现差异显著。
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models
- 针对巴斯克、印地、斯瓦希里语设计针对性句法测试
- 部分句法任务模型表现良好,但涉及间接宾语和介词短语的搭配任务困难
- 发现多语言BERT存在时态偏见,XGLM-4.5B性能低于同类模型
语言模型具备习得类人句法知识的能力。已有研究通过定向句法评估测试衡量其在英语等高资源语言中的句法泛化能力,但对低资源语言中模型句法泛化能力的理解仍不充分,而这些语言承载了全球大部分句法多样性。本研究为巴斯克语、印地语和斯瓦希里语开发了定向句法评估测试,并用于评估五种开源多语言Transformer语言模型。结果表明,部分句法任务对模型而言相对容易,但涉及巴斯克语间接宾语一致性和斯瓦希里语介词短语跨结构一致性的任务则极具挑战。此外,还发现了公开多语言Transformer模型的问题:多语言BERT在印地语中表现出对习惯体的偏差,XGLM-4.5B在性能上逊于相似规模的其他模型。
原文摘要 · Abstract (English)
Language models (LMs) are capable of acquiring elements of human-like syntactic knowledge. Targeted syntactic evaluation tests have been employed to measure how well they form generalizations about syntactic phenomena in high-resource languages such as English. However, we still lack a thorough understanding of LMs' capacity for syntactic generalizations in low-resource languages, which are responsible for much of the diversity of syntactic patterns worldwide. In this study, we develop targeted syntactic evaluation tests for three low-resource languages (Basque, Hindi, and Swahili) and use them to evaluate five families of open-access multilingual Transformer LMs. We find that some syntactic tasks prove relatively easy for LMs while others (agreement in sentences containing indirect objects in Basque, agreement across a prepositional phrase in Swahili) are challenging. We additionally uncover issues with publicly available Transformers, including a bias toward the habitual aspect in Hindi in multilingual BERT and underperformance compared to similar-sized models in XGLM-4.5B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。