用高级提示工程提升大模型在科研知识图谱中的属性抽取准确率
Evaluating improvements on using Large Language Models (LLMs) for property extraction in the Open Research Knowledge Graph (ORKG)
- 采用先进提示工程优化大模型的属性抽取能力
- 提升结果与现有知识图谱属性的匹配率,增强一致性
- 适合关注科研知识图谱构建与数据标准化的研究者
当前研究显示大语言模型(LLMs)在构建学术知识图谱(SKGs)方面潜力巨大,但关系抽取仍具挑战性。本研究基于前人对GPT-3.5、Llama 2和Mistral等模型在科学文献属性抽取中的评估,发现其表现中等,需微调以更好契合科研任务。本研究进一步评估了高级提示工程的效果,结果表明其能显著提升性能。同时,研究将属性抽取扩展至与ORKG已有属性匹配,通过API检索实现。实验显示,经提示工程优化的结果与ORKG属性的匹配比例更高,增强了模型输出与知识图谱的一致性。该方法有助于解决属性不一致问题,通过赋予唯一URI和标准化术语,符合链接数据与FAIR原则。这显著提升了ORKG内容在后续研究比较等任务中的可用性。研究最后提出未来改进方向。
原文摘要 · Abstract (English)
Current research highlights the great potential of Large Language Models (LLMs) for constructing Scholarly Knowledge Graphs (SKGs). One particularly complex step in this process is relation extraction, aimed at identifying suitable properties to describe the content of research. This study builds directly on previous research of three Open Research Knowledge Graph (ORKG) team members who assessed the readiness of LLMs such as GPT-3.5, Llama 2, and Mistral for property extraction in scientific literature. Given the moderate performance observed, the previous work concluded that fine-tuning is needed to improve these models' alignment with scientific tasks and their emulation of human expertise. Expanding on this prior experiment, this study evaluates the impact of advanced prompt engineering techniques and demonstrates that these techniques can highly significantly enhance the results. Additionally, this study extends the property extraction process to include property matching to existing ORKG properties, which are retrieved via the API. The evaluation reveals that results generated through advanced prompt engineering achieve a higher proportion of matches with ORKG properties, further emphasizing the enhanced alignment achieved. Moreover, this lays the groundwork for addressing challenges such as the inconsistency of ORKG properties, an issue highlighted in prior studies. By assigning unique URIs and using standardized terminology, this work increases the consistency of the properties, fulfilling a crucial aspect of Linked Data and FAIR principles - core commitments of ORKG. This, in turn, significantly enhances the applicability of ORKG content for subsequent tasks such as comparisons of research publications. Finally, the study concludes with recommendations for future improvements in the overall property extraction process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。