arXiv:2506.23051cs.CL2025-06被引 1

首个面向20世纪初巴西葡语的历史文本命名实体识别数据集

MariNER: A Dataset for Historical Brazilian Portuguese Named Entity Recognition

  • 构建了超过9000句手工标注的早期巴西葡语文本数据集
  • 首次为历史巴西葡语提供高质量金标准,支持数字人文研究
  • 适合研究历史文本、语言演变与低资源语言NLP的学者

命名实体识别(NER)是自然语言处理中识别和分类文本中实体提及的基础任务。尽管英语等语言拥有大量高质量资源,但巴西葡语在该任务上仍缺乏足够数量的金标准数据集,尤其在特定领域。本文关注数字人文背景下历史文本的NER需求。为此,本文构建了MariNER:Mapeamento e Anotações de Registros hIstóricos para NER(历史记录映射与标注用于NER),这是首个针对20世纪初巴西葡语的金标准数据集,包含超过9,000句手工标注句子。同时,我们评估并比较了当前主流NER模型在该数据集上的表现。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) is a fundamental Natural Language Processing (NLP) task that aims to identify and classify entity mentions in texts across different categories. While languages such as English possess a large number of high-quality resources for this task, Brazilian Portuguese still lacks in quantity of gold-standard NER datasets, especially when considering specific domains. Particularly, this paper considers the importance of NER for analyzing historical texts in the context of digital humanities. To address this gap, this work outlines the construction of MariNER: \textit{Mapeamento e Anotações de Registros hIstóricos para NER} (Mapping and Annotation of Historical Records for NER), the first gold-standard dataset for early 20th-century Brazilian Portuguese, with more than 9,000 manually annotated sentences. We also assess and compare the performance of state-of-the-art NER models for the dataset.

命名实体识别历史文本巴西葡语数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。