arXiv:2410.17355cs.CL2024-10被引 3

发现大模型对罕见实体识别能力差,提出新方法评估长尾问题。

All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing

  • 基于启发式方法估算预训练数据中实体分布,定位长尾问题。
  • 实验证明仅依赖模型参数的方案在稀有实体上性能显著下降。
  • 建议融合外部知识提升对少见实体的识别效果,适合长尾场景研究者。

由于能够从大规模语料中获取世界知识,预训练语言模型(PLMs)被广泛用于超细粒度实体标注任务,该任务标签空间极大。本文提出一种新颖启发式方法,在无法获取预训练数据的情况下近似实体的预训练分布。系统性地证明,仅依赖PLM参数化知识的方法在预训练分布长尾部分的实体上表现显著不足,而引入外部知识的方法可部分缓解此问题。研究结果表明,需超越纯模型依赖的方案,才能有效处理低频实体。

原文摘要 · Abstract (English)

Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity typing tasks where the space of labels is extremely large. In this work, we explore the limitations of the knowledge acquired by PLMs by proposing a novel heuristic to approximate the pre-training distribution of entities when the pre-training data is unknown. Then, we systematically demonstrate that entity-typing approaches that rely solely on the parametric knowledge of PLMs struggle significantly with entities at the long tail of the pre-training distribution, and that knowledge-infused approaches can account for some of these shortcomings. Our findings suggest that we need to go beyond PLMs to produce solutions that perform well for infrequent entities.

实体识别长尾问题预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。