arXiv:2502.14561cs.CLcs.DL2025-02中稿 · publication on TPD…被引 9

用少量数据让通用大模型学会预测引用意图,效果媲美专业模型。

Can LLMs Predict Citation Intent? An Experimental Analysis of In-context Learning and Fine-tuning on Open LLMs

  • 用提示学习和微调让通用大模型理解论文引用目的。
  • 在SciCite上提升8%的F1分数,ACL-ARC上提升4.3%。
  • 适合想低成本部署引用分析系统的研究人员使用。

本研究探讨开放型大语言模型通过提示学习与微调预测引用意图的能力。不同于依赖领域预训练模型(如SciBERT)的传统方法,我们证明通用大模型仅需少量任务特定数据即可适配该任务。在五个主流开源大模型家族中评估了十二种模型变体,采用零样本、单样本、少样本及多样本提示。通过大规模提示学习实验,确定了表现最佳的模型与提示参数。随后通过微调该模型,相较指令微调基线,在SciCite数据集上实现相对F1分数提升8%,在ACL-ARC数据集上提升4.3%。研究为模型选择与提示工程提供了重要参考。此外,我们公开了端到端评估框架与模型,供未来研究使用。

原文摘要 · Abstract (English)

This work investigates the ability of open Large Language Models (LLMs) to predict citation intent through in-context learning and fine-tuning. Unlike traditional approaches relying on domain-specific pre-trained models like SciBERT, we demonstrate that general-purpose LLMs can be adapted to this task with minimal task-specific data. We evaluate twelve model variations across five prominent open LLM families using zero-, one-, few-, and many-shot prompting. Our experimental study identifies the top-performing model and prompting parameters through extensive in-context learning experiments. We then demonstrate the significant impact of task-specific adaptation by fine-tuning this model, achieving a relative F1-score improvement of 8% on the SciCite dataset and 4.3% on the ACL-ARC dataset compared to the instruction-tuned baseline. These findings provide valuable insights for model selection and prompt engineering. Additionally, we make our end-to-end evaluation framework and models openly available for future use.

引用预测提示学习大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。