用大模型加人工反馈,自动提取材料科学中的结构化知识。
LLMs4SchemaDiscovery: A Human-in-the-Loop Workflow for Scientific Schema Mining with Large Language Models
- 大模型结合人类反馈,迭代优化文本中的属性组织。
- 在原子层沉积领域生成语义丰富的结构化数据模式。
- 适合需要高质量领域知识的科研人员和数据构建者。
从非结构化文本中提取结构化信息对建模现实世界过程至关重要,但传统模式挖掘依赖半结构化数据,限制了可扩展性。本文提出schema-miner,一种结合大语言模型与人工反馈的新工具,用于自动化并优化模式提取。通过迭代工作流程,该工具从文本中组织属性,融入专家输入,并整合领域特定本体以增强语义深度。应用于材料科学——特别是原子层沉积领域,结果表明,经专家引导的大模型能生成适用于多种实际应用的语义丰富模式。
原文摘要 · Abstract (English)
Extracting structured information from unstructured text is crucial for modeling real-world processes, but traditional schema mining relies on semi-structured data, limiting scalability. This paper introduces schema-miner, a novel tool that combines large language models with human feedback to automate and refine schema extraction. Through an iterative workflow, it organizes properties from text, incorporates expert input, and integrates domain-specific ontologies for semantic depth. Applied to materials science--specifically atomic layer deposition--schema-miner demonstrates that expert-guided LLMs generate semantically rich schemas suitable for diverse real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。