构建首个巴斯克-西班牙语混用语料库,助力语言混合处理研究
EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
- 从现有语料中用语言识别模型筛选混用文本,人工验证确保质量
- 形成首个自然来源的巴斯克-西班牙语混用语料库,支持相关研究
- 适合研究多语种交互、语言接触及低资源语言处理的学者使用
代码切换(CS)仍是自然语言处理中的重大挑战,主要因缺乏相关数据。在伊比利亚半岛北部巴斯克语与西班牙语接触背景下,混用现象在正式与非正式交流中频繁出现。然而,用于分析该现象并支持模型开发与评估的资源几乎空白。本文提出首个自然来源的巴斯克-西班牙语代码切换语料库构建方法:利用语言识别模型从已有语料中识别混用文本,并经人工验证获得可靠实例集。我们介绍该语料库的特性,并以EuskañolDS为名公开发布。
原文摘要 · Abstract (English)
Code-switching (CS) remains a significant challenge in Natural Language Processing (NLP), mainly due a lack of relevant data. In the context of the contact between the Basque and Spanish languages in the north of the Iberian Peninsula, CS frequently occurs in both formal and informal spontaneous interactions. However, resources to analyse this phenomenon and support the development and evaluation of models capable of understanding and generating code-switched language for this language pair are almost non-existent. We introduce a first approach to develop a naturally sourced corpus for Basque-Spanish code-switching. Our methodology consists of identifying CS texts from previously available corpora using language identification models, which are then manually validated to obtain a reliable subset of CS instances. We present the properties of our corpus and make it available under the name EuskañolDS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。