arXiv:2410.01946cs.CL2024-10EMNLP被引 5

自动增强提示词,提升科学文献细粒度分类效果

SciPrompt: Knowledge-augmented Prompting for Fine-grained Categorization of Scientific Topics

  • 用语义相关术语自动构建提示词库,替代人工设计
  • 在少样本和零样本下优于现有方法,尤其擅长新兴领域分类
  • 适合需要快速适配新科学领域的文本分类任务

基于提示的微调已成为从预训练语言模型中提取信息的关键方法,广泛应用于文本分类等任务。在低资源场景下,提示微调已达到与全量微调相当的性能。以往研究依赖人工设计的提示模板和标签映射器(verbalizer),将标签空间映射到类别空间,以掩码语言建模方式解决分类问题。然而,跨领域、细粒度的提示微调仍缺乏自动化的标签术语扩充方法,主要受限于手动选择领域标签术语的成本高、需专业知识。为此,我们提出SciPrompt框架,可自动检索科学文献中与主题相关的术语,用于增强提示词。我们引入一种新策略,利用相关性得分作为权重,在模型微调时提升预测性能。实验表明,该方法在少样本和零样本设置下显著优于当前最优提示微调方法,尤其在细粒度及新兴科学主题分类任务中表现突出。

原文摘要 · Abstract (English)

Prompt-based fine-tuning has become an essential method for eliciting information encoded in pre-trained language models for a variety of tasks, including text classification. For multi-class classification tasks, prompt-based fine-tuning under low-resource scenarios has resulted in performance levels comparable to those of fully fine-tuning methods. Previous studies have used crafted prompt templates and verbalizers, mapping from the label terms space to the class space, to solve the classification problem as a masked language modeling task. However, cross-domain and fine-grained prompt-based fine-tuning with an automatically enriched verbalizer remains unexplored, mainly due to the difficulty and costs of manually selecting domain label terms for the verbalizer, which requires humans with domain expertise. To address this challenge, we introduce SciPrompt, a framework designed to automatically retrieve scientific topic-related terms for low-resource text classification tasks. To this end, we select semantically correlated and domain-specific label terms within the context of scientific literature for verbalizer augmentation. Furthermore, we propose a new verbalization strategy that uses correlation scores as additional weights to enhance the prediction performance of the language model during model tuning. Our method outperforms state-of-the-art, prompt-based fine-tuning methods on scientific text classification tasks under few and zero-shot settings, especially in classifying fine-grained and emerging scientific topics.

提示工程科学分类少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。