arXiv:2409.00369cs.CL2024-09被引 77

测试GPT-4在信息抽取上的表现,发现与顶尖方法差距明显

An Empirical Study on Information Extraction using Large Language Models

  • 用提示工程改进GPT-4的信息抽取能力
  • GPT-4在抽取任务上仍落后于专业模型
  • 简单提示策略可推广至其他LLM和任务

大型语言模型(如OpenAI的GPT系列)在自然语言处理任务中展现出强大潜力,因此被广泛尝试用于信息抽取(IE)。为评估当前最先进的模型GPT-4在信息抽取方面的表现,本文从性能、评价标准、鲁棒性及错误类型四个维度进行实证分析。结果表明,GPT-4与现有最先进(SOTA)IE方法之间存在显著性能差距。为此,基于大模型的人类化特征,本文提出并分析了一系列简单有效的提示工程方法,具有良好的泛化性,可应用于其他大模型和自然语言处理任务。大量实验验证了这些方法的有效性,同时也揭示了其在提升抽取能力方面仍存在的局限。

原文摘要 · Abstract (English)

Human-like large language models (LLMs), especially the most powerful and popular ones in OpenAI's GPT family, have proven to be very helpful for many natural language processing (NLP) related tasks. Therefore, various attempts have been made to apply LLMs to information extraction (IE), which is a fundamental NLP task that involves extracting information from unstructured plain text. To demonstrate the latest representative progress in LLMs' information extraction ability, we assess the information extraction ability of GPT-4 (the latest version of GPT at the time of writing this paper) from four perspectives: Performance, Evaluation Criteria, Robustness, and Error Types. Our results suggest a visible performance gap between GPT-4 and state-of-the-art (SOTA) IE methods. To alleviate this problem, considering the LLMs' human-like characteristics, we propose and analyze the effects of a series of simple prompt-based methods, which can be generalized to other LLMs and NLP tasks. Rich experiments show our methods' effectiveness and some of their remaining issues in improving GPT-4's information extraction ability.

信息抽取提示工程GPT-4

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。