arXiv:2412.00426cs.CL2024-12被引 1

用少量标注数据+大量无标签数据,实现跨领域命名实体识别新突破

Few-Shot Domain Adaptation for Named-Entity Recognition via Joint Constrained k-Means and Subspace Selection

  • 融合约束k均值与子空间选择,联合优化模型在少样本下的迁移能力
  • 在多个英文数据集上达到当前最佳效果,仅需少量标注即有效
  • 适合需要快速适配新领域但标注资源稀缺的NLP应用

命名实体识别(NER)通常需要大量标注数据,限制了其在实体定义各异的新领域的应用。本文针对少样本NER问题,提出一种弱监督算法,结合少量标注数据与大量无标签数据进行知识迁移。所提方法在传统k均值基础上引入标签监督、聚类规模约束以及领域特异性判别子空间选择,构建统一框架,在多个英文数据集上实现了少样本NER的当前最优性能。

原文摘要 · Abstract (English)

Named-entity recognition (NER) is a task that typically requires large annotated datasets, which limits its applicability across domains with varying entity definitions. This paper addresses few-shot NER, aiming to transfer knowledge to new domains with minimal supervision. Unlike previous approaches that rely solely on limited annotated data, we propose a weakly supervised algorithm that combines small labeled datasets with large amounts of unlabeled data. Our method extends the k-means algorithm with label supervision, cluster size constraints and domain-specific discriminative subspace selection. This unified framework achieves state-of-the-art results in few-shot NER on several English datasets.

少样本学习命名实体识别领域自适应无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。