16种非洲语言的NLI模型表现不随数据量单调提升,需针对性优化。
Sample-Size Scaling of the African Languages NLI Evaluation

- 在16种非洲语言上测试小样本(50-500)下多语言模型性能
- 多数语言出现早期饱和或性能下降,低资源下方差大
- 提示需为每种语言定制数据集与模型策略
非洲语言标注数据极少,其标注量是否能稳定提升下游任务性能尚不明确。本研究基于AfriXNLI基准,对16种非洲语言开展系统性样本量缩放实验。在受控条件下,使用约0.6B参数的XLM-R Large和AfroXLM-R Large两个多语言Transformer模型,在50至500个标注样本范围内进行微调,并对随机子采样结果取平均。结果显示,与通常认为性能随数据增加而单调上升的观点相反,性能呈现强烈语言依赖且常非单调变化。部分语言在小样本时即出现性能饱和或下降,低资源环境下波动显著。这表明仅增加数据量不足以保证非洲语言自然语言推理任务的稳定收益,亟需构建更具语言敏感性的数据集并发展更强的多语言建模方法。
原文摘要 · Abstract (English)
African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance. The study is a systematic sample-size scaling study of natural language inference (NLI) on 16 African languages based on the AfriXNLI benchmark. Under controlled conditions, two multilingual transformer models with roughly 0.6B parameters XLM-R Large fine-tuned on XNLI and AfroXLM-R Large are tested on sample sizes of between 50 and 500 labeled examples and average their results across random subsampling runs. As opposed to the usual belief of monotonic increase with increased data, we find a strongly language sensitive and often non-monotonic scaling behavior. Some languages show early saturation or decrease in performance with sample size as well as high variance in low resource regimes. These results indicate that the volume of data is not enough to guarantee stable profits to African NLI, creating the necessity of language sensitive datasets creation and stronger multi-lingual modelling strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。