评测9种模型在6种印地语系语言中的语言特性编码能力与鲁棒性。
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
- 构建包含4.7万句的多语言基准集IndicSentEval,覆盖13种文本扰动。
- 通用模型在英语中编码一致,但在印地语系语言中表现参差,特定模型更优。
- 通用模型对删除名词/动词等扰动更具鲁棒性,适合跨语言任务研究。
基于Transformer的模型彻底改变了自然语言处理领域。为理解其高性能原因并评估可靠性,已有研究关注模型编码语言特性(如句法、语义)的程度及其在输入扰动下的鲁棒性。但这些研究主要集中于BERT和英语。本文首次系统考察了9种多语言Transformer模型(7种通用、2种印地语专用)在6种印地语系语言中对8类语言特性的编码能力与13种扰动下的鲁棒性。为此,我们构建了全新的多语言基准数据集IndicSentEval,含约47,000条句子。探针分析显示:尽管所有模型在英语中表现一致,但在印地语系语言中结果不一;印地语专用模型在本地语言上表现更好;有趣的是,通用模型在删除名词/动词等扰动下总体更鲁棒。本研究揭示了主流多语言Transformer模型在不同印地语系语言中的优劣,代码与数据已公开。
原文摘要 · Abstract (English)
Transformer-based models have revolutionized the field of natural language processing. To understand why they perform so well and to assess their reliability, several studies have focused on questions such as: Which linguistic properties are encoded by these models, and to what extent? How robust are these models in encoding linguistic properties when faced with perturbations in the input text? However, these studies have mainly focused on BERT and the English language. In this paper, we investigate similar questions regarding encoding capability and robustness for 8 linguistic properties across 13 different perturbations in 6 Indic languages, using 9 multilingual Transformer models (7 universal and 2 Indic-specific). To conduct this study, we introduce a novel multilingual benchmark dataset, IndicSentEval, containing approximately $\sim$47K sentences. Surprisingly, our probing analysis of surface, syntactic, and semantic properties reveals that while almost all multilingual models demonstrate consistent encoding performance for English, they show mixed results for Indic languages. As expected, Indic-specific multilingual models capture linguistic properties in Indic languages better than universal models. Intriguingly, universal models broadly exhibit better robustness compared to Indic-specific models, particularly under perturbations such as dropping both nouns and verbs, dropping only verbs, or keeping only nouns. Overall, this study provides valuable insights into probing and perturbation-specific strengths and weaknesses of popular multilingual Transformer-based models for different Indic languages. We make our code and dataset publicly available [https://github.com/aforakhilesh/IndicBertology].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。