arXiv:2409.13266cs.CL2024-09EMNLP被引 7

利用引用上下文自动生成关键词,无需人工标注即可提升跨领域关键词生成效果。

Unsupervised Domain Adaptation for Keyphrase Generation using Citation Contexts

  • 从引用文本中挖掘高质量伪标签作为合成训练数据
  • 在三个不同领域均显著超越强基线模型
  • 适合缺乏标注数据的学术文本领域迁移任务

将关键词生成模型迁移到新领域通常需要少量有标注的领域内数据进行微调。然而,为文档标注关键词成本高昂且不切实际,需依赖专业标注人员。本文提出Silk,一种无监督方法,通过从引用上下文中提取银标准关键词,构建合成标注数据以实现领域适应。在三个不同领域的大量实验表明,该方法生成的合成样本质量高,显著且一致地提升了领域内性能,优于多个强基线。

原文摘要 · Abstract (English)

Adapting keyphrase generation models to new domains typically involves few-shot fine-tuning with in-domain labeled data. However, annotating documents with keyphrases is often prohibitively expensive and impractical, requiring expert annotators. This paper presents silk, an unsupervised method designed to address this issue by extracting silver-standard keyphrases from citation contexts to create synthetic labeled data for domain adaptation. Extensive experiments across three distinct domains demonstrate that our method yields high-quality synthetic samples, resulting in significant and consistent improvements in in-domain performance over strong baselines.

关键词生成无监督学习领域自适应引用分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。