arXiv:2603.02150cs.CLcs.AI2026-03被引 2

构建首个通用犯罪实体识别数据集CrimeNER-db,助力执法信息提取。

Named-Entity Recognition in the Crime Domain (CrimeNER): Case Study and Dataset

  • 定义4类粗粒度、21类细粒度犯罪实体,覆盖恐怖袭击与司法通报场景。
  • 含1500+标注文档,验证了模型在零样本与少样本下的泛化能力。
  • 适合法律科技、情报分析与信息抽取方向的研究者使用。

从犯罪相关文档中提取关键信息对执法机构至关重要,可视为命名实体识别(NER)任务。然而,真实世界犯罪场景的高质量标注数据严重不足。为此,我们提出CrimeNER,一项犯罪领域NER的案例研究,并构建了包含超过1500份标注文档的通用犯罪领域命名实体识别数据库(CrimeNER-db),数据源自公开的恐怖袭击报告及美国司法部新闻稿。我们定义了4种粗粒度犯罪实体类型和21种细粒度实体类型。通过全监督微调模型以及零样本和少样本实验,评估了数据库的质量与模型泛化能力。该数据集已开源于GitHub。

原文摘要 · Abstract (English)

The extraction of critical information from crime-related documents is a crucial task for law enforcement agencies. The extraction of this information can be interpreted as a Named-Entity Recognition (NER) task. However, there is a considerable lack of adequately annotated data on general real-world crime scenarios. To address this issue, we present CrimeNER, a case study of crime-related NER, and a general crime-related Named-Entity Recognition database (CrimeNER-db), consisting of more than 1.5K annotated documents extracted from public reports of terrorist attacks and the US Department of Justice's press notes. We define 4 coarse types of crime entity and 21 fine-grained entity types. We address the quality of the presented database with experiments using fully supervised finetuned general NER models and zero- and few-shot experiments to address the generalization capabilities. The database is available on GitHub.

命名实体识别犯罪分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。