arXiv:2504.04963cs.CL2025-04

提出新方法提升远距离监督命名实体识别的准确率

Constraint Multi-class Positive and Unlabeled Learning for Distantly Supervised Named Entity Recognition

  • 引入约束因子优化多类别正负样本学习的风险估计
  • 在两个基准数据集上显著降低误漏率,优于现有方法
  • 适合处理标注不全的自动标注场景

远距离监督命名实体识别(DS-NER)通过外部知识库自动生成训练标签,避免人工标注。但因知识库不完整,常导致高误漏率。本文提出约束多类别正负样本学习(CMPU)方法,在多类正样本风险估计中引入约束因子,使模型对少量正样本更具鲁棒性,不易过拟合。理论分析证明了该方法的有效性。在两个基于不同外部知识源构建的基准数据集上的实验表明,CMPU显著优于现有DS-NER方法。

原文摘要 · Abstract (English)

Distantly supervised named entity recognition (DS-NER) has been proposed to exploit the automatically labeled training data by external knowledge bases instead of human annotations. However, it tends to suffer from a high false negative rate due to the inherent incompleteness. To address this issue, we present a novel approach called \textbf{C}onstraint \textbf{M}ulti-class \textbf{P}ositive and \textbf{U}nlabeled Learning (CMPU), which introduces a constraint factor on the risk estimator of multiple positive classes. It suggests that the constraint non-negative risk estimator is more robust against overfitting than previous PU learning methods with limited positive data. Solid theoretical analysis on CMPU is provided to prove the validity of our approach. Extensive experiments on two benchmark datasets that were labeled using diverse external knowledge sources serve to demonstrate the superior performance of CMPU in comparison to existing DS-NER methods.

命名实体识别远距离监督正负样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。