arXiv:2503.14173cs.CL2025-03

针对加泰罗尼亚语低资源现状,通过领域标注数据微调模型提升命名实体识别效果。

NERCat: Fine-Tuning for Enhanced Named Entity Recognition in Catalan

  • 用加泰罗尼亚语电视字幕人工标注数据微调GLiNER模型。
  • 在法律、产品、设施等类别上F1值显著提升,尤其改善了低频实体识别。
  • 适合关注小语种NLP、领域定制化模型的研究者与开发者。

命名实体识别(NER)是自然语言处理中从非结构化文本中提取结构化信息的关键环节。然而,对于加泰罗尼亚语这类低资源语言,由于高质量标注数据匮乏,NER系统性能常受限。本文提出NERCat,即针对加泰罗尼亚语优化的GLiNER[1]模型微调版本。我们使用人工标注的加泰罗尼亚语电视字幕数据集进行训练与微调,重点覆盖政治、体育和文化等领域的实体。评估结果显示,精度、召回率和F1分数均有显著提升,尤其在法律、产品、设施等代表性不足的实体类别上表现突出。本研究证明了领域特定微调在低资源语言中的有效性,并强调通过人工标注构建高质量数据集对提升加泰罗尼亚语NLP应用的潜力。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) is a critical component of Natural Language Processing (NLP) for extracting structured information from unstructured text. However, for low-resource languages like Catalan, the performance of NER systems often suffers due to the lack of high-quality annotated datasets. This paper introduces NERCat, a fine-tuned version of the GLiNER[1] model, designed to improve NER performance specifically for Catalan text. We used a dataset of manually annotated Catalan television transcriptions to train and fine-tune the model, focusing on domains such as politics, sports, and culture. The evaluation results show significant improvements in precision, recall, and F1-score, particularly for underrepresented named entity categories such as Law, Product, and Facility. This study demonstrates the effectiveness of domain-specific fine-tuning in low-resource languages and highlights the potential for enhancing Catalan NLP applications through manual annotation and high-quality datasets.

命名实体识别小语种NLP微调加泰罗尼亚语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。