arXiv:2411.00046cs.CLcs.AI2024-11被引 14

用大模型辅助生物数据整理,让专家更快更准地建库。

CurateGPT: A flexible language-model assisted biocuration tool

  • 用大模型代理自动查文献、对齐术语、整合外部数据。
  • 相比直接问大模型,能链接真实数据源并验证每条结论。
  • 适合生物信息学、知识库建设者快速处理海量文献。

高效的数据驱动生物发现依赖于数据整理:一项耗时的任务,包括查找、组织、提炼、整合、解释、注释和验证多样信息,使其成为数据库与知识库可用的结构化形式。准确高效的数字资产整理对确保数据符合可查找、可访问、可互操作、可重用(FAIR)标准、可信且可持续至关重要。然而,专家整理者面临时间和资源严重不足的问题。每天大量新论文发表,远超人工整理能力。生成式AI,特别是经过指令微调的大语言模型(LLMs),为辅助人类整理提供了新可能。我们的设计哲学是将生成式AI的新能力与更精确的方法结合。整理者的任务可通过智能体实现推理、搜索本体、跨外部源整合知识,这些原本需要大量手动工作。我们开发的基于大模型的注释工具CurateGPT,融合了生成式AI与可信知识库及文献源的力量。CurateGPT简化了整理流程,提升了常见工作流中的协作与效率。相比直接使用大模型,CurateGPT的智能体能访问模型训练数据之外的信息,并为每条结论提供直接的数据来源链接。这帮助整理者、研究者和工程师扩大整理规模,跟上科学数据不断增长的步伐。

原文摘要 · Abstract (English)

Effective data-driven biomedical discovery requires data curation: a time-consuming process of finding, organizing, distilling, integrating, interpreting, annotating, and validating diverse information into a structured form suitable for databases and knowledge bases. Accurate and efficient curation of these digital assets is critical to ensuring that they are FAIR, trustworthy, and sustainable. Unfortunately, expert curators face significant time and resource constraints. The rapid pace of new information being published daily is exceeding their capacity for curation. Generative AI, exemplified by instruction-tuned large language models (LLMs), has opened up new possibilities for assisting human-driven curation. The design philosophy of agents combines the emerging abilities of generative AI with more precise methods. A curator's tasks can be aided by agents for performing reasoning, searching ontologies, and integrating knowledge across external sources, all efforts otherwise requiring extensive manual effort. Our LLM-driven annotation tool, CurateGPT, melds the power of generative AI together with trusted knowledge bases and literature sources. CurateGPT streamlines the curation process, enhancing collaboration and efficiency in common workflows. Compared to direct interaction with an LLM, CurateGPT's agents enable access to information beyond that in the LLM's training data and they provide direct links to the data supporting each claim. This helps curators, researchers, and engineers scale up curation efforts to keep pace with the ever-increasing volume of scientific data.

生物信息大模型数据整理知识库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。