arXiv:2603.25189cs.CL2026-03

整理巴斯克语方言资源,助力方言自然语言处理研究

A Catalog of Basque Dialectal Resources: Online Collections and Standard-to-Dialectal Adaptations

  • 分类整理在线方言文本与标准语转方言数据
  • 手动构建三类方言的高质量评估数据集
  • 验证自动转换数据可作为银标准替代方案

近年来,方言自然语言处理研究受限于数据稀缺。本文系统梳理了当前可用的巴斯克语方言数据与资源,分为两类:一是直接源自方言的在线文本(如新闻、广播、推文、词典、语法、视频等),二是从标准语转为方言的数据。对于手动转换,将XNLI自然语言推理数据集的测试集手工转换为西巴斯克、中巴斯克和纳瓦拉-拉普尔迪安三种方言,形成高质量平行金标准评估数据集。对于自动转换,对自动转换的物理常识数据集(BasPhyCowest)进行母语者人工评估,以判断其是否可作为完整人工标注的可行替代(即银标准数据)。结果表明,自动转换数据具备一定可用性,可减少人工成本。

原文摘要 · Abstract (English)

Recent research on dialectal NLP has identified data scarcity as a primary limitation. To address this limitation, this paper presents a catalog of contemporary Basque dialectal data and resources, offering a systematic and comprehensive compilation of the dialectal data currently available in Basque. Two types of data sources have been distinguished: online data originally written in some dialect, and standard-to-dialect adapted data. The former includes all dialectal data that can be found online, such as news and radio sites, informal tweets, as well as online resources such as dictionaries, atlases, grammar rules, or videos. The latter consists of data that has been adapted from the standard variety to dialectal varieties, either manually or automatically. Regarding the manual adaptation, the test split of the XNLI Natural Language Inference dataset was manually adapted into three Basque dialects: Western, Central, and Navarrese-Lapurdian, yielding a high-quality parallel gold standard evaluation dataset. With respect to the automatic dialectal adaptation, the automatically adapted physical commonsense dataset (BasPhyCowest) underwent additional manual evaluation by native speakers to assess its quality and determine whether it could serve as a viable substitute for full manual adaptation (i.e., silver data creation).

方言处理数据资源巴斯克语自动转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。