arXiv:2604.19505cs.IRcs.CL2026-04

融合论文亮点内容,显著提升无监督关键词提取效果

Enhancing Unsupervised Keyword Extraction in Academic Papers through Integrating Highlights with Abstract

  • 将论文亮点与摘要结合,作为关键词提取输入
  • 在四个数据集上,融合方法性能优于单一来源
  • 适合需要高精度关键词抽取的研究者使用

学术论文的自动关键词提取是自然语言处理与信息检索的重要方向。以往研究多基于摘要和参考文献,本文聚焦于论文亮点部分——一个概括核心发现与贡献的简短描述,能为读者提供快速概览。我们观察到亮点中蕴含丰富关键词信息,可有效补充摘要。为探究融合亮点对无监督关键词提取的影响,我们在四个无监督模型上评估了三种输入场景:仅用摘要、仅用亮点、以及两者结合。实验在计算机科学(CS)与图书馆与信息科学(LIS)数据集上进行,结果表明,将摘要与亮点结合能显著提升提取性能。此外,我们分析了摘要与亮点在关键词覆盖范围和内容上的差异,探讨其对提取结果的影响。相关数据与代码已公开于 https://github.com/xiangyi-njust/Highlight-KPE。

原文摘要 · Abstract (English)

Automatic keyword extraction from academic papers is a key area of interest in natural language processing and information retrieval. Although previous research has mainly focused on utilizing abstract and references for keyword extraction, this paper focuses on the highlights section - a summary describing the key findings and contributions, offering readers a quick overview of the research. Our observations indicate that highlights contain valuable keyword information that can effectively complement the abstract. To investigate the impact of incorporating highlights into unsupervised keyword extraction, we evaluate three input scenarios: using only the abstract, the highlights, and a combination of both. Experiments conducted with four unsupervised models on Computer Science (CS), Library and Information Science (LIS) datasets reveal that integrating the abstract with highlights significantly improves extraction performance. Furthermore, we examine the differences in keyword coverage and content between abstract and highlights, exploring how these variations influence extraction outcomes. The data and code are available at https://github.com/xiangyi-njust/Highlight-KPE.

关键词提取无监督学习论文分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。