arXiv:2410.18040cs.CLcs.AI2024-10被引 2

用提示词让大模型生成俄语科研摘要关键词,效果优于传统方法。

Key Algorithms for Keyphrase Generation: Instruction-Based LLMs for Russian Scientific Keyphrases

  • 用文本提示代替微调,直接生成俄语关键词。
  • 零样本和少样本提示方法均超越传统基线模型。
  • 适合需要快速部署、数据稀缺的俄语NLP任务。

关键词提取是自然语言处理中的挑战性任务,应用广泛。由于俄语形态丰富且训练数据有限,现有监督与无监督方法面临诸多限制。近期针对英文文本的研究表明,大语言模型(LLMs)可通过文本提示实现无需特定任务微调的关键词生成,取得优异效果。本文评估了基于提示的方法在生成俄语科学摘要关键词上的表现。首先对比零样本、少样本提示方法、微调模型与无监督方法的性能;其次分析少样本设置下关键词示例的选择策略。通过人工评估生成关键词并由专家分析模型优劣。结果表明,仅使用简单文本提示的提示方法即可超越常见基线,展现出良好潜力。

原文摘要 · Abstract (English)

Keyphrase selection is a challenging task in natural language processing that has a wide range of applications. Adapting existing supervised and unsupervised solutions for the Russian language faces several limitations due to the rich morphology of Russian and the limited number of training datasets available. Recent studies conducted on English texts show that large language models (LLMs) successfully address the task of generating keyphrases. LLMs allow achieving impressive results without task-specific fine-tuning, using text prompts instead. In this work, we access the performance of prompt-based methods for generating keyphrases for Russian scientific abstracts. First, we compare the performance of zero-shot and few-shot prompt-based methods, fine-tuned models, and unsupervised methods. Then we assess strategies for selecting keyphrase examples in a few-shot setting. We present the outcomes of human evaluation of the generated keyphrases and analyze the strengths and weaknesses of the models through expert assessment. Our results suggest that prompt-based methods can outperform common baselines even using simple text prompts.

关键词生成大模型俄语NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。