利用大模型数据训练信息抽取模型,无需人工标注。
Cuckoo: An IE Free Rider Hatched by Massive Nutrition in LLM's Nest
- 将语言模型的预测任务转为已存在文本的抽取任务
- 用1.026亿条数据训练出性能超越现有模型的Cuckoo
- 能自动跟随大模型升级,无需额外标注
大量高质量数据,包括预训练文本和后训练标注,被精心准备以孵化先进的大语言模型(LLM)。相比之下,信息抽取(IE)的预训练数据(如BIO标注序列)难以规模化。我们表明,通过将下一个词的预测重构为上下文已有词的抽取,信息抽取模型可作为大模型资源的免费搭乘者。具体而言,所提出的下一词抽取(NTE)范式,训练出一个名为Cuckoo的通用信息抽取模型,使用1.026亿条从大模型预训练和后训练数据中转换的抽取数据。在少样本设置下,Cuckoo能有效适应传统和复杂指令遵循型信息抽取任务,表现优于现有预训练信息抽取模型。作为免费搭乘者,Cuckoo可自然随大模型数据准备的持续进展而进化,无需额外人工努力即可受益于大模型训练流程的改进。
原文摘要 · Abstract (English)
Massive high-quality data, both pre-training raw texts and post-training annotations, have been carefully prepared to incubate advanced large language models (LLMs). In contrast, for information extraction (IE), pre-training data, such as BIO-tagged sequences, are hard to scale up. We show that IE models can act as free riders on LLM resources by reframing next-token \emph{prediction} into \emph{extraction} for tokens already present in the context. Specifically, our proposed next tokens extraction (NTE) paradigm learns a versatile IE model, \emph{Cuckoo}, with 102.6M extractive data converted from LLM's pre-training and post-training data. Under the few-shot setting, Cuckoo adapts effectively to traditional and complex instruction-following IE with better performance than existing pre-trained IE models. As a free rider, Cuckoo can naturally evolve with the ongoing advancements in LLM data preparation, benefiting from improvements in LLM training pipelines without additional manual effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。