为10种低资源语言构建了词义消歧语料,助力跨语言迁移研究。
SenWiCh: Sense-Annotation of Low-Resource Languages for WiC using Hybrid Methods
- 采用半自动方法高效标注多义词语境数据
- 在10种低资源语言上验证了跨语言迁移效果
- 适合关注多语言NLP与公平评估的研究者
本文针对低资源语言缺乏高质量评测数据的问题,提出一种半自动标注方法,构建了涵盖十种低资源语言的多义词语境(WiC)标注数据集,覆盖多种语言家族和书写系统。通过在这些语言上开展基于WiC格式的实验,验证了数据集的有效性。结果表明,在低资源环境下,针对性的数据构建对多义词消歧和跨语言迁移至关重要。所发布的数据集与代码将支持更公平、稳健且真正的多语言自然语言处理研究。
原文摘要 · Abstract (English)
This paper addresses the critical need for high-quality evaluation datasets in low-resource languages to advance cross-lingual transfer. While cross-lingual transfer offers a key strategy for leveraging multilingual pretraining to expand language technologies to understudied and typologically diverse languages, its effectiveness is dependent on quality and suitable benchmarks. We release new sense-annotated datasets of sentences containing polysemous words, spanning ten low-resource languages across diverse language families and scripts. To facilitate dataset creation, the paper presents a demonstrably beneficial semi-automatic annotation method. The utility of the datasets is demonstrated through Word-in-Context (WiC) formatted experiments that evaluate transfer on these low-resource languages. Results highlight the importance of targeted dataset creation and evaluation for effective polysemy disambiguation in low-resource settings and transfer studies. The released datasets and code aim to support further research into fair, robust, and truly multilingual NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。