arXiv:2411.17863cs.CLcs.AI2024-11中稿 · presentation at th…被引 4

解决长文档关键词提取难题,提升大段文本理解能力

LongKey: Keyphrase Extraction for Long Documents

  • 采用编码器模型捕捉长文本语义,增强关键词候选表示
  • 在LDKP及六个新数据集上表现超越现有方法
  • 适合处理学术论文、报告等超长文档的关键词抽取

信息爆炸时代,人工标注海量文档日益不切实际。自动化关键词提取可识别文本中的代表性术语。然而,现有方法多聚焦短文档(不超过512个词元),难以处理长上下文文档。本文提出LongKey框架,利用基于编码器的语言模型捕获长文档的复杂语义,并通过最大池化嵌入器增强关键词候选表征。在全面的LDKP数据集及六个不同领域的新数据集上验证,LongKey持续优于现有无监督与基于语言模型的方法。结果表明,LongKey具备强泛化能力与卓越性能,推动了跨文本长度与领域的关键词提取技术发展。

原文摘要 · Abstract (English)

In an era of information overload, manually annotating the vast and growing corpus of documents and scholarly papers is increasingly impractical. Automated keyphrase extraction addresses this challenge by identifying representative terms within texts. However, most existing methods focus on short documents (up to 512 tokens), leaving a gap in processing long-context documents. In this paper, we introduce LongKey, a novel framework for extracting keyphrases from lengthy documents, which uses an encoder-based language model to capture extended text intricacies. LongKey uses a max-pooling embedder to enhance keyphrase candidate representation. Validated on the comprehensive LDKP datasets and six diverse, unseen datasets, LongKey consistently outperforms existing unsupervised and language model-based keyphrase extraction methods. Our findings demonstrate LongKey's versatility and superior performance, marking an advancement in keyphrase extraction for varied text lengths and domains.

关键词提取长文本语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。