小模型微调后精准构建航天任务知识图谱,效果优于大模型。
Autoregressive Language Models for Knowledge Base Population: A case study in the space mission domain
- 用领域语料微调自回归模型,实现端到端知识库填充。
- 小模型在航天任务数据上准确率超越更大模型。
- 无需在提示中包含本体,节省上下文空间,适合部署。
知识库填充(KBP)在组织中通过利用领域文本资料,对知识库进行动态维护和更新具有关键作用。受大型语言模型日益增长的上下文窗口支持启发,我们提出对自回归语言模型进行微调,以实现端到端的KBP。本案例研究聚焦于航天任务知识图谱的构建。为训练模型,我们利用现有领域资源生成了一个用于端到端KBP的数据集。结果表明,经过微调的小型模型在KBP任务中可达到与更大模型相当甚至更高的准确率。这些专用于KBP的小模型具备低成本部署和推理的优势。此外,它们无需在提示中包含本体信息,从而为额外输入文本或输出序列化释放更多上下文空间。
原文摘要 · Abstract (English)
Knowledge base population KBP plays a crucial role in populating and maintaining knowledge bases up-to-date in organizations by leveraging domain corpora. Motivated by the increasingly large context windows supported by large language models, we propose to fine-tune an autoregressive language model for end-toend KPB. Our case study involves the population of a space mission knowledge graph. To fine-tune the model we generate a dataset for end-to-end KBP tapping into existing domain resources. Our case study shows that fine-tuned language models of limited size can achieve competitive and even higher accuracy than larger models in the KBP task. Smaller models specialized for KBP offer affordable deployment and lower-cost inference. Moreover, KBP specialist models do not require the ontology to be included in the prompt, allowing for more space in the context for additional input text or output serialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。