用大模型和概念图谱发现材料科学中未被察觉的新兴研究方向。
Predicting New Research Directions in Materials Science using Large Language Models and Concept Graphs
- 通过大模型提取论文摘要中的核心概念,构建科学文献的概念图谱。
- 基于历史数据预测概念组合,识别潜在的新研究方向,性能优于传统方法。
- 专家访谈验证其启发性,适合希望开拓创新思路的研究者使用。
由于发表的研究论文呈指数增长,科学家难以阅读本领域全部文献。本文探索利用大语言模型(LLMs)从材料科学领域的论文摘要中提取核心概念与语义信息,发现人类未注意到的关联,从而提出近中期有潜力的研究方向。实验表明,LLMs比自动关键词提取方法更高效地构建概念图谱,作为科学文献的抽象表示。我们训练了一个机器学习模型,基于历史数据预测新兴的概念组合,即新研究思路。结果表明,融入语义概念信息可显著提升预测性能。通过与领域专家的定性访谈验证了模型的应用价值:它能通过预测尚未被研究的课题组合,激发材料科学家的创造性思考。
原文摘要 · Abstract (English)
Due to an exponential increase in published research articles, it is impossible for individual scientists to read all publications, even within their own research field. In this work, we investigate the use of large language models (LLMs) for the purpose of extracting the main concepts and semantic information from scientific abstracts in the domain of materials science to find links that were not noticed by humans and thus to suggest inspiring near/mid-term future research directions. We show that LLMs can extract concepts more efficiently than automated keyword extraction methods to build a concept graph as an abstraction of the scientific literature. A machine learning model is trained to predict emerging combinations of concepts, i.e. new research ideas, based on historical data. We demonstrate that integrating semantic concept information leads to an increased prediction performance. The applicability of our model is demonstrated in qualitative interviews with domain experts based on individualized model suggestions. We show that the model can inspire materials scientists in their creative thinking process by predicting innovative combinations of topics that have not yet been investigated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。