arXiv:2511.20143cs.CLcs.AI2025-11

用图像增强方法提升网格模型识别断续实体能力

SEDA: A Self-Adapted Entity-Centric Data Augmentation for Boosting Gird-based Discontinuous NER Models

  • 借鉴图像增强思想,对网格模型进行自适应数据扩增
  • 在多个数据集上使断续实体识别F1提升3.7%-8.4%
  • 适合需要精准识别跨句断续实体的研究者

命名实体识别(NER)是自然语言处理中的关键任务,但对断续实体的识别仍具挑战。主要难点在于文本分段:传统方法常误分或完全遗漏跨句断续实体,严重影响识别准确率。为此,本文针对分段与遗漏问题提出解决方案。近期研究表明,网格标注方法因灵活的标注机制和鲁棒架构,在信息抽取中表现优异。基于此,本文将图像数据增强技术(如裁剪、缩放、填充)引入网格模型,以提升其对断续实体的识别能力与分段适应性。实验表明,传统分段方法难以捕捉跨句断续实体,导致性能下降;而本文提出的增强型网格模型取得显著改进。在CADEC、ShARe13和ShARe14数据集上的评估显示,整体F1分数提升1-2.5%,断续实体部分提升达3.7%-8.4%,充分验证了该方法的有效性。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) is a critical task in natural language processing, yet it remains particularly challenging for discontinuous entities. The primary difficulty lies in text segmentation, as traditional methods often missegment or entirely miss cross-sentence discontinuous entities, significantly affecting recognition accuracy. Therefore, we aim to address the segmentation and omission issues associated with such entities. Recent studies have shown that grid-tagging methods are effective for information extraction due to their flexible tagging schemes and robust architectures. Building on this, we integrate image data augmentation techniques, such as cropping, scaling, and padding, into grid-based models to enhance their ability to recognize discontinuous entities and handle segmentation challenges. Experimental results demonstrate that traditional segmentation methods often fail to capture cross-sentence discontinuous entities, leading to decreased performance. In contrast, our augmented grid models achieve notable improvements. Evaluations on the CADEC, ShARe13, and ShARe14 datasets show F1 score gains of 1-2.5% overall and 3.7-8.4% for discontinuous entities, confirming the effectiveness of our approach.

NER断续实体数据增强网格模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。