arXiv:2411.05057cs.CL2024-11被引 6

梳理命名实体识别三十年技术演进,从监督到无监督学习。

A Brief History of Named Entity Recognition

  • 回顾1996年以来NER技术从监督学习到无监督学习的演变
  • 系统对比各类方法在信息抽取、语义标注等任务中的表现
  • 适合对自然语言处理发展史感兴趣的读者

当今世界大量信息存储于知识库中。命名实体识别(NER)是从原始文本中提取、消歧和链接实体以构建结构化知识库的过程。具体而言,它旨在识别并分类文本中对信息抽取、语义标注、问答系统、本体构建等至关重要的实体。自1996年首次出现以来,NER技术在过去三十年中不断演进。本文综述了用于NER的技术发展历程,并从监督学习到新兴的无监督学习方法进行了系统比较。

原文摘要 · Abstract (English)

A large amount of information in today's world is now stored in knowledge bases. Named Entity Recognition (NER) is a process of extracting, disambiguation, and linking an entity from raw text to insightful and structured knowledge bases. More concretely, it is identifying and classifying entities in the text that are crucial for Information Extraction, Semantic Annotation, Question Answering, Ontology Population, and so on. The process of NER has evolved in the last three decades since it first appeared in 1996. In this survey, we study the evolution of techniques employed for NER and compare the results, starting from supervised to the developing unsupervised learning methods.

命名实体识别自然语言处理技术综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。