arXiv:2512.10004cs.AIcs.CL2025-12中稿 · AAAI被引 9

SciEx框架提升科学文献细粒度信息提取效率与灵活性。

Exploring LLMs for Scientific Information Extraction Using The SciEx Framework

  • 模块化设计分离解析、检索、抽取与聚合环节
  • 在三个领域数据集上实现准确一致的细粒度信息提取
  • 支持灵活接入新模型与推理策略,适应快速变化的格式需求

大语言模型(LLMs)被广泛视为自动化科学信息提取的强大工具。然而,现有方法和工具在应对科学文献的现实挑战时表现不佳:长上下文文档、多模态内容,以及将多篇出版物中的细粒度信息统一为标准化格式时存在不一致问题。当目标数据模式或抽取本体快速变化时,重构或微调现有系统尤为困难。我们提出SciEx,一个模块化且可组合的框架,将PDF解析、多模态检索、信息抽取与聚合等关键组件解耦。该设计简化了按需数据提取流程,同时支持扩展性与新模型、提示策略和推理机制的灵活集成。我们在涵盖三个科学主题的数据集上评估了SciEx在准确性和一致性提取细粒度信息方面的能力。研究结果为当前基于LLM的流水线提供了实用洞见,揭示其优势与局限。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly touted as powerful tools for automating scientific information extraction. However, existing methods and tools often struggle with the realities of scientific literature: long-context documents, multi-modal content, and reconciling varied and inconsistent fine-grained information across multiple publications into standardized formats. These challenges are further compounded when the desired data schema or extraction ontology changes rapidly, making it difficult to re-architect or fine-tune existing systems. We present SciEx, a modular and composable framework that decouples key components including PDF parsing, multi-modal retrieval, extraction, and aggregation. This design streamlines on-demand data extraction while enabling extensibility and flexible integration of new models, prompting strategies, and reasoning mechanisms. We evaluate SciEx on datasets spanning three scientific topics for its ability to extract fine-grained information accurately and consistently. Our findings provide practical insights into both the strengths and limitations of current LLM-based pipelines.

信息提取科学文献LLM框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。