arXiv:2409.01087cs.CLcs.AI2024-09综述被引 17

综述预训练模型在关键词提取与生成中的应用

Pre-Trained Language Models for Keyphrase Prediction: A Review

  • 按任务类型分类,系统梳理预训练模型在关键词提取与生成中的方法
  • 涵盖监督、无监督等多类训练范式,分析其性能差异
  • 适合从事NLP文本摘要与信息抽取研究的读者

关键词预测(KP)对于识别文档中能概括内容的关键短语至关重要。近年来,自然语言处理技术发展出基于深度学习的高效KP模型。然而,现有研究对预训练语言模型在关键词提取(KPE)与生成(KPG)两类任务中的联合探索仍显不足,形成文献空白。本文全面回顾预训练语言模型用于关键词预测(PLM-KP)的研究进展,这些模型通过监督、无监督、半监督及自监督等多种学习方式在大规模文本语料上进行训练。文章为PLM-KPE和PLM-KPG建立合适的分类体系,深入分析两类任务的特性与挑战,并指出未来在关键词预测方向的潜在研究路径。

原文摘要 · Abstract (English)

Keyphrase Prediction (KP) is essential for identifying keyphrases in a document that can summarize its content. However, recent Natural Language Processing (NLP) advances have developed more efficient KP models using deep learning techniques. The limitation of a comprehensive exploration jointly both keyphrase extraction and generation using pre-trained language models spotlights a critical gap in the literature, compelling our survey paper to bridge this deficiency and offer a unified and in-depth analysis to address limitations in previous surveys. This paper extensively examines the topic of pre-trained language models for keyphrase prediction (PLM-KP), which are trained on large text corpora via different learning (supervisor, unsupervised, semi-supervised, and self-supervised) techniques, to provide respective insights into these two types of tasks in NLP, precisely, Keyphrase Extraction (KPE) and Keyphrase Generation (KPG). We introduce appropriate taxonomies for PLM-KPE and KPG to highlight these two main tasks of NLP. Moreover, we point out some promising future directions for predicting keyphrases.

关键词预测预训练模型NLP综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。