arXiv:2504.18142cs.CLcs.AI2025-04被引 3

首个面向乌尔都语教育文本的命名实体识别数据集

EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)

  • 构建乌尔都语教育领域手动标注数据集
  • 覆盖13类教育相关实体,解决标注资源匮乏问题
  • 针对乌尔都语语法复杂性提出标注方案

命名实体识别(NER)在自然语言处理中至关重要,用于从非结构化文本中识别并分类人名、组织、地点、日期等实体。尽管高资源语言和通用领域已有大量研究,但乌尔都语在教育等特定领域的NER仍严重缺乏。这主要由于缺乏教育类文本的标注数据,导致现有模型难以准确识别学术角色、课程名称和机构术语。本文首次构建了教育领域乌尔都语文本的标注数据集EDU-NER-2025,包含13类关键教育实体。详细说明了标注流程与规范,并分析了正式乌尔都语文本中的形态复杂性和歧义性等语言挑战。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) plays a pivotal role in various Natural Language Processing (NLP) tasks by identifying and classifying named entities (NEs) from unstructured data into predefined categories such as person, organization, location, date, and time. While extensive research exists for high-resource languages and general domains, NER in Urdu particularly within domain-specific contexts like education remains significantly underexplored. This is Due to lack of annotated datasets for educational content which limits the ability of existing models to accurately identify entities such as academic roles, course names, and institutional terms, underscoring the urgent need for targeted resources in this domain. To the best of our knowledge, no dataset exists in the domain of the Urdu language for this purpose. To achieve this objective this study makes three key contributions. Firstly, we created a manually annotated dataset in the education domain, named EDU-NER-2025, which contains 13 unique most important entities related to education domain. Second, we describe our annotation process and guidelines in detail and discuss the challenges of labelling EDU-NER-2025 dataset. Third, we addressed and analyzed key linguistic challenges, such as morphological complexity and ambiguity, which are prevalent in formal Urdu texts.

命名实体识别乌尔都语教育NLP数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。