arXiv:2508.06877cs.CLcs.AI2025-08

自动对齐多源命名实体数据集标签,提升低资源领域识别效果

ESNERA: Empirical and semantic named entity alignment for named entity dataset merging

  • 结合统计与语义相似性,自动对齐不同数据集的实体标签
  • 在金融领域小样本数据上实现性能提升,融合后模型准确率更高
  • 方法可解释、易扩展,适合多领域文本标注数据整合

命名实体识别(NER)是自然语言处理的基础任务,广泛应用于多个领域。尽管深度学习显著提升了识别性能,但其依赖大规模高质量标注数据,而构建这些数据成本高、耗时长,成为研究瓶颈。现有数据集合并方法多依赖人工标签映射或构建标签图,缺乏可解释性和可扩展性。为此,本文提出一种基于标签相似性的自动对齐方法,融合经验与语义相似性,并采用贪心成对合并策略统一不同数据集的标签空间。实验分为两阶段:首先将三个现有NER数据集合并为统一语料库,对原有性能影响最小;其次将该语料库与一个自建的小规模金融领域数据集融合。结果表明,该方法有效实现了数据集合并,并在低资源金融领域提升了NER性能。本研究提供了一种高效、可解释且可扩展的多源NER语料整合方案。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) is a fundamental task in natural language processing. It remains a research hotspot due to its wide applicability across domains. Although recent advances in deep learning have significantly improved NER performance, they rely heavily on large, high-quality annotated datasets. However, building these datasets is expensive and time-consuming, posing a major bottleneck for further research. Current dataset merging approaches mainly focus on strategies like manual label mapping or constructing label graphs, which lack interpretability and scalability. To address this, we propose an automatic label alignment method based on label similarity. The method combines empirical and semantic similarities, using a greedy pairwise merging strategy to unify label spaces across different datasets. Experiments are conducted in two stages: first, merging three existing NER datasets into a unified corpus with minimal impact on NER performance; second, integrating this corpus with a small-scale, self-built dataset in the financial domain. The results show that our method enables effective dataset merging and enhances NER performance in the low-resource financial domain. This study presents an efficient, interpretable, and scalable solution for integrating multi-source NER corpora.

命名实体识别数据融合低资源学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。