arXiv:2411.03039cs.CLcs.IR2024-11中稿 · JCDL 2024被引 1

用共享关键词合并文档,自动生成高质量训练数据

Self-Compositional Data Augmentation for Scientific Keyphrase Generation

  • 基于共现关键词相似性合并文档生成合成样本
  • 在三个领域数据集上均提升生成效果,计算机领域更具代表性
  • 无需外部数据,适合资源有限的科研场景

当前先进关键短语生成模型需大量标注数据才能表现良好,但获取带关键短语标注的文档既困难又昂贵。为此,本文提出一种自组合数据增强方法:根据训练文档的共享关键词衡量其相关性,并将相似文档组合生成合成样本。该方法优势在于不依赖外部数据或资源,即可生成保持领域一致性的额外训练样本。在涵盖三个不同领域的多个数据集上的实验表明,该方法能持续提升关键短语生成性能。对计算机科学领域的定性分析进一步验证了生成结果在代表性上的改进。

原文摘要 · Abstract (English)

State-of-the-art models for keyphrase generation require large amounts of training data to achieve good performance. However, obtaining keyphrase-labeled documents can be challenging and costly. To address this issue, we present a self-compositional data augmentation method. More specifically, we measure the relatedness of training documents based on their shared keyphrases, and combine similar documents to generate synthetic samples. The advantage of our method lies in its ability to create additional training samples that keep domain coherence, without relying on external data or resources. Our results on multiple datasets spanning three different domains, demonstrate that our method consistently improves keyphrase generation. A qualitative analysis of the generated keyphrases for the Computer Science domain confirms this improvement towards their representativity property.

关键短语生成数据增强自组合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。