arXiv:2506.17580cs.IRcs.AI2025-06

用智能流程从科学文献中精准提取深度知识,效率提升超80%

Context-Aware Scientific Knowledge Extraction on Linked Open Data using Large Language Models

  • 构建树状结构工作流,动态筛选与问题相关的上下文信息
  • 实验显示处理文本量减少80%以上,召回率显著高于基线方法
  • 适合科研人员快速整合跨领域科学知识,尤其在药物研发等场景

科学文献的爆炸式增长使研究人员难以高效提取和整合知识。传统搜索引擎返回大量来源但缺乏直接答案,通用大模型虽简洁却常缺深度或滞后信息。具备搜索能力的大模型受限于上下文窗口,输出短且不完整。本文提出WISE(智能科学知识提取工作流),通过结构化流程系统性地提取、精炼并排序查询相关知识。WISE采用基于大模型的树形架构,聚焦于与查询对齐、上下文感知且无冗余的信息。通过动态评分与排名机制优先呈现各来源的独特贡献,自适应停止条件降低计算开销。实验在与HBB基因相关疾病上验证,WISE将处理文本量减少超过80%,召回率显著优于搜索引擎及其它基于LLM的方法。ROUGE与BLEU指标表明其输出更具独特性,新提出的层级评估指标显示其提供更深入的信息。此外,该工作流可拓展至药物发现、材料科学与社会科学等领域,实现从非结构化论文与网络资源中高效提取与合成知识。

原文摘要 · Abstract (English)

The exponential growth of scientific literature challenges researchers extracting and synthesizing knowledge. Traditional search engines return many sources without direct, detailed answers, while general-purpose LLMs may offer concise responses that lack depth or omit current information. LLMs with search capabilities are also limited by context window, yielding short, incomplete answers. This paper introduces WISE (Workflow for Intelligent Scientific Knowledge Extraction), a system addressing these limits by using a structured workflow to extract, refine, and rank query-specific knowledge. WISE uses an LLM-powered, tree-based architecture to refine data, focusing on query-aligned, context-aware, and non-redundant information. Dynamic scoring and ranking prioritize unique contributions from each source, and adaptive stopping criteria minimize processing overhead. WISE delivers detailed, organized answers by systematically exploring and synthesizing knowledge from diverse sources. Experiments on HBB gene-associated diseases demonstrate WISE reduces processed text by over 80% while achieving significantly higher recall over baselines like search engines and other LLM-based approaches. ROUGE and BLEU metrics reveal WISE's output is more unique than other systems, and a novel level-based metric shows it provides more in-depth information. We also explore how the WISE workflow can be adapted for diverse domains like drug discovery, material science, and social science, enabling efficient knowledge extraction and synthesis from unstructured scientific papers and web sources.

知识提取大模型应用科学文献智能工作流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。