arXiv:2411.03484cond-mat.mtrl-scics.IR2024-11被引 11

用大模型自动提取文献中的多孔材料合成信息,免调优且防幻觉。

Automated, LLM enabled extraction of synthesis details for reticular materials from scientific literature

  • 基于提示工程与上下文学习,无需训练即可分类和提取文献段落
  • 仅需少量示例即在五类开源大模型上表现优异
  • 适合材料科研人员快速获取合成工艺,降低人工标注负担

从科学文献中自动化提取知识可加速材料发现。本文研究了利用大语言模型(LLMs)提取多孔材料合成方案的方法。为此,我们提出知识提取流程(KEP),实现大模型辅助的段落分类与信息抽取自动化。通过在一组开源大模型上应用提示工程与上下文学习(ICL),证明大模型可在不进行微调或训练的情况下,从PDF文档中准确提取化学信息,且幻觉风险较低。在五类开源大模型的段落分类与信息提取任务中,即使仅在ICL提示中包含少量示例,模型性能依然出色。结果表明,该KEP方法具有显著降低人工标注与数据整理工作量的潜力。

原文摘要 · Abstract (English)

Automated knowledge extraction from scientific literature can potentially accelerate materials discovery. We have investigated an approach for extracting synthesis protocols for reticular materials from scientific literature using large language models (LLMs). To that end, we introduce a Knowledge Extraction Pipeline (KEP) that automatizes LLM-assisted paragraph classification and information extraction. By applying prompt engineering with in-context learning (ICL) to a set of open-source LLMs, we demonstrate that LLMs can retrieve chemical information from PDF documents, without the need for fine-tuning or training and at a reduced risk of hallucination. By comparing the performance of five open-source families of LLMs in both paragraph classification and information extraction tasks, we observe excellent model performance even if only few example paragraphs are included in the ICL prompts. The results show the potential of the KEP approach for reducing human annotations and data curation efforts in automated scientific knowledge extraction.

知识提取大模型材料发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。