arXiv:2503.00597cs.CLcs.AI2025-03NAACL被引 1

用提示工程提升大模型零样本关键词生成能力

Zero-Shot Keyphrase Generation: Investigating Specialized Instructions and Multi-Sample Aggregation on Large Language Models

  • 设计针对性提示指令,引导大模型聚焦关键词任务
  • 多样本聚合策略使生成结果准确率显著提升
  • 适用于无标注数据场景下的自动摘要与信息抽取

关键词是概括文档主题的核心短语。关键词生成是自然语言处理中长期存在的任务,旨在为给定文档自动生成关键词。尽管该任务过去已通过多种模型广泛研究,但仅有少数工作对大语言模型(LLMs)在此任务上的表现进行了初步分析。鉴于大语言模型在自然语言处理领域的重大影响,有必要对其在关键词生成中的潜力进行更深入的考察。本文致力于满足这一需求,重点研究开源指令微调模型(Phi-3、Llama-3)和闭源模型GPT-4o在零样本场景下的表现。系统评估了在提示中加入任务相关专业指令的效果,并设计了针对该任务的自洽式多样本聚合策略,实验表明所提方法显著优于基线。

原文摘要 · Abstract (English)

Keyphrases are the essential topical phrases that summarize a document. Keyphrase generation is a long-standing NLP task for automatically generating keyphrases for a given document. While the task has been comprehensively explored in the past via various models, only a few works perform some preliminary analysis of Large Language Models (LLMs) for the task. Given the impact of LLMs in the field of NLP, it is important to conduct a more thorough examination of their potential for keyphrase generation. In this paper, we attempt to meet this demand with our research agenda. Specifically, we focus on the zero-shot capabilities of open-source instruction-tuned LLMs (Phi-3, Llama-3) and the closed-source GPT-4o for this task. We systematically investigate the effect of providing task-relevant specialized instructions in the prompt. Moreover, we design task-specific counterparts to self-consistency-style strategies for LLMs and show significant benefits from our proposals over the baselines.

关键词生成大模型零样本提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。