arXiv:2410.13318cs.CLcs.AI2024-10被引 2

首个阿拉伯语-英语混用命名实体识别数据集与方法

Computational Approaches to Arabic-English Code-Switching

  • 构建首个阿拉伯语-英语混用命名实体标注语料库
  • 通过上下文嵌入与数据增强使识别准确率提升
  • 提出词内语言识别方法,助力混合文本分析

自然语言处理(NLP)是解决语言处理、分析与生成的重要计算方法,广泛应用于自动校正、语音识别等场景。尽管英语相关研究丰富,但对现代标准阿拉伯语和方言阿拉伯语的关注较少。全球化推动了语言混用(Code-Switching, CS)现象,尤其在埃及社交网络中,阿拉伯语与英语频繁交替使用,形成大量混合文本。现有研究尚未针对阿拉伯语-英语混用数据开展核心NLP任务。本文聚焦命名实体识别(NER),首次构建了该领域的标注语料库,并应用前沿技术改进模型性能。通过引入混用上下文嵌入与数据增强策略,显著提升了NER识别效果。此外,提出多种词内语言识别方法,以判断混合文本的语言归属及实体属性。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech recognition. While significant research has focused on NLP tasks for the English language, less attention has been given to Modern Standard Arabic and Dialectal Arabic. Globalization has also contributed to the rise of Code-Switching (CS), where speakers mix languages within conversations and even within individual words (intra-word CS). This is especially common in Arab countries, where people often switch between dialects or between dialects and a foreign language they master. CS between Arabic and English is frequent in Egypt, especially on social media. Consequently, a significant amount of code-switched content can be found online. Such code-switched data needs to be investigated and analyzed for several NLP tasks to tackle the challenges of this multilingual phenomenon and Arabic language challenges. No work has been done before for several integral NLP tasks on Arabic-English CS data. In this work, we focus on the Named Entity Recognition (NER) task and other tasks that help propose a solution for the NER task on CS data, e.g., Language Identification. This work addresses this gap by proposing and applying state-of-the-art techniques for Modern Standard Arabic and Arabic-English NER. We have created the first annotated CS Arabic-English corpus for the NER task. Also, we apply two enhancement techniques to improve the NER tagger on CS data using CS contextual embeddings and data augmentation techniques. All methods showed improvements in the performance of the NER taggers on CS data. Finally, we propose several intra-word language identification approaches to determine the language type of a mixed text and identify whether it is a named entity or not.

命名实体识别语言混用阿拉伯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。