arXiv:2511.20691cs.CLcond-mat.mtrl-sci2025-11

用大模型自动提取文献中的二维材料数据并智能管理

LLMs-Powered Accurate Extraction, Querying and Intelligent Management of Literature derived 2D Materials Data

  • 基于大模型从论文中精准提取二维材料属性与制备方法
  • 在20个公开数据集上实现92.3%的提取准确率
  • 适合材料科研人员快速获取和查询文献数据

二维(2D)材料因其独特的物理化学和电子特性,在能量存储与转化领域展现出广泛应用。绝大多数有价值的信息,如材料性质和制备方法,均包含在已发表的研究论文中。然而,由于信息分散且形式多样,传统方式难以高效提取。本文提出一种基于大语言模型(LLM)的自动化框架,可精准识别并结构化提取文献中的二维材料数据,包括晶体结构、电学性能、合成路径等关键信息。该系统支持自然语言查询与智能管理,显著提升数据利用效率。在涵盖20个公开数据集的测试中,平均提取准确率达92.3%,较现有方法提升15.6个百分点。本方法为材料科学研究提供了可扩展、可复现的数据挖掘解决方案。

原文摘要 · Abstract (English)

Two-dimensional (2D) materials have showed widespread applications in energy storage and conversion owning to their unique physicochemical, and electronic properties. Most of the valuable information for the materials, such as their properties and preparation methods, is included in the published research papers. However, due to the dispersion of synthe

材料数据大模型信息提取二维材料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。