arXiv:2411.12000cs.CLcs.AI2024-11被引 3

用自动微调大模型从科学文献中提取结构化数据,提升知识转化效率。

ByteScience: Bridging Unstructured Scientific Literature and Structured Data with Auto Fine-tuned Large Language Model in Token Granularity

  • 基于DARWIN模型在标记数据极少时实现高精度提取
  • 依托AWS搭建自动化流程,支持用户自定义模型开发
  • 适合科研人员快速构建领域知识库,推动自然信息学发展

自然语言处理(NLP)广泛用于从长文本中提取结构化信息,但因科学文本的领域特性、复杂预处理及多层级设备级信息粒度,仍面临挑战。为此,我们提出ByteScience——一个非营利性的基于云的自动微调大语言模型平台,旨在从海量科学文献中提取结构化科学数据并合成新知识。该平台基于开源的科学专用微调大模型DARWIN,部署于亚马逊云服务(AWS),提供自动化、易用的定制模型开发与数据提取工作流。仅需少量高质量标注文章,即可实现显著精度。该工具有效缩短了科学文献到结构化知识的转化路径,助力自然信息学进步。

原文摘要 · Abstract (English)

Natural Language Processing (NLP) is widely used to supply summarization ability from long context to structured information. However, extracting structured knowledge from scientific text by NLP models remains a challenge because of its domain-specific nature to complex data preprocessing and the granularity of multi-layered device-level information. To address this, we introduce ByteScience, a non-profit cloud-based auto fine-tuned Large Language Model (LLM) platform, which is designed to extract structured scientific data and synthesize new scientific knowledge from vast scientific corpora. The platform capitalizes on DARWIN, an open-source, fine-tuned LLM dedicated to natural science. The platform was built on Amazon Web Services (AWS) and provides an automated, user-friendly workflow for custom model development and data extraction. The platform achieves remarkable accuracy with only a small amount of well-annotated articles. This innovative tool streamlines the transition from the science literature to structured knowledge and data and benefits the advancements in natural informatics.

科学文本挖掘大模型微调结构化数据提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。