分析50篇论文,揭示关键词生成领域的数据、评估与模型问题。
An Analysis of Datasets, Metrics and Models in Keyphrase Generation
- 系统梳理50篇论文,总结关键词生成方法演进
- 发现基准数据集高度相似,评估指标导致性能虚高
- 开源强模型助力后续研究,弥补预训练资源不足
关键词生成旨在从文档中自动生成能概括内容的词或短语。近年来,该领域在模型架构、数据资源和应用场景等方面持续发展,但整体进展缺乏系统性回顾与分析。本文通过分析超过50篇相关研究,全面梳理了该任务的最新进展、局限与开放挑战。研究发现当前评估实践存在严重问题:常用基准数据集之间相似性过高,且评估指标计算不一致,导致性能结果被显著夸大。此外,为缓解预训练模型稀缺的问题,我们发布了一个基于PLM的强大关键词生成模型,以推动未来研究发展。
原文摘要 · Abstract (English)
Keyphrase generation refers to the task of producing a set of words or phrases that summarises the content of a document. Continuous efforts have been dedicated to this task over the past few years, spreading across multiple lines of research, such as model architectures, data resources, and use-case scenarios. Yet, the current state of keyphrase generation remains unknown as there has been no attempt to review and analyse previous work. In this paper, we bridge this gap by presenting an analysis of over 50 research papers on keyphrase generation, offering a comprehensive overview of recent progress, limitations, and open challenges. Our findings highlight several critical issues in current evaluation practices, such as the concerning similarity among commonly-used benchmark datasets and inconsistencies in metric calculations leading to overestimated performances. Additionally, we address the limited availability of pre-trained models by releasing a strong PLM-based model for keyphrase generation as an effort to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。