arXiv:2410.23478cs.CLcs.HC2024-10被引 1

Collage让非NLP研究者快速测试科学文献信息抽取模型

Collage: Decomposable Rapid Prototyping for Information Extraction on Scientific PDFs

  • 支持HuggingFace、LLM等多模型快速接入与对比
  • 可处理科学PDF,提供中间状态可视化调试
  • 适合材料科学等领域的文献综述自动化

近年来,领域特定的信息抽取工具和多模态预训练变压器模型持续发展。尽管科学家能更清晰地评估和应用这些系统,但模型难以比较:输入格式各异,常为黑箱且缺乏故障分析,且很少支持科学出版物最常见的PDF格式。本文提出Collage,一个用于科学PDF上信息抽取的快速原型设计、可视化与评估工具。Collage可直接使用HuggingFace分词器、多种大语言模型及其他任务特定模型,并提供可扩展接口以加速新模型实验。此外,通过展示处理过程中的中间状态,帮助开发者和用户检查、调试并理解建模流程。我们在材料科学文献综述场景中展示了该系统的应用。

原文摘要 · Abstract (English)

Recent years in NLP have seen the continued development of domain-specific information extraction tools for scientific documents, alongside the release of increasingly multimodal pretrained transformer models. While the opportunity for scientists outside of NLP to evaluate and apply such systems to their own domains has never been clearer, these models are difficult to compare: they accept different input formats, are often black-box and give little insight into processing failures, and rarely handle PDF documents, the most common format of scientific publication. In this work, we present Collage, a tool designed for rapid prototyping, visualization, and evaluation of different information extraction models on scientific PDFs. Collage allows the use and evaluation of any HuggingFace token classifier, several LLMs, and multiple other task-specific models out of the box, and provides extensible software interfaces to accelerate experimentation with new models. Further, we enable both developers and users of NLP-based tools to inspect, debug, and better understand modeling pipelines by providing granular views of intermediate states of processing. We demonstrate our system in the context of information extraction to assist with literature review in materials science.

信息抽取科学文献快速原型PDF处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。