arXiv:2412.19191q-bio.BMcs.AI2024-12EMNLP被引 8

首个面向多组学序列的大规模指令微调数据集,提升大模型生物理解能力。

Biology-Instructions: A Dataset and Benchmark for Multi-Omics Sequence Understanding Capability of Large Language Models

  • 构建包含DNA/RNA/蛋白质的多组学序列指令数据集
  • 现有大模型在未微调下对多组学任务表现不佳
  • 提出三阶段训练框架,显著增强模型生物学推理能力

大语言模型在通用领域表现出色,但在多组学生物学应用中仍鲜有探索。为此,我们推出Biology-Instructions,首个面向多组学生物序列(包括DNA、RNA、蛋白质及多分子)的大规模指令微调数据集。该数据集连接大语言模型与复杂生物序列任务,提升其泛化性与推理能力,同时保持对话流畅性。我们还揭示了当前顶尖大模型在缺乏专门训练时,在多组学任务上的显著局限。为克服此问题,我们提出ChatMultiOmics,一种基于新型三阶段训练流程的强基线模型,通过Biology-Instructions验证其在多组学理解上的优越性能。两项资源均公开可用,推动大模型在多组学分析中的深度融合。Biology-Instructions可访问:https://github.com/hhnqqq/Biology-Instructions。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown remarkable capabilities in general domains, but their application to multi-omics biology remains underexplored. To address this gap, we introduce Biology-Instructions, the first large-scale instruction-tuning dataset for multi-omics biological sequences, including DNA, RNA, proteins, and multi-molecules. This dataset bridges LLMs and complex biological sequence-related tasks, enhancing their versatility and reasoning while maintaining conversational fluency. We also highlight significant limitations of current state-of-the-art LLMs on multi-omics tasks without specialized training. To overcome this, we propose ChatMultiOmics, a strong baseline with a novel three-stage training pipeline, demonstrating superior biological understanding through Biology-Instructions. Both resources are publicly available, paving the way for better integration of LLMs in multi-omics analysis. The Biology-Instructions is publicly available at: https://github.com/hhnqqq/Biology-Instructions.

多组学大模型生物序列指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。