将基因组数据与大模型结合,实现可解释的生物推理
BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM Model
- 用大模型直接解析基因序列,进行多步逻辑推理
- 疾病通路预测准确率提升至98%,变异影响预测平均提高15%
- 能解释未知生物实体的推理过程,适合生物科研人员使用
从复杂的基因组数据中获取深层且可解释的生物学推理仍是人工智能在生物学领域的主要挑战。现有DNA基础模型虽擅长序列表征,但在多步推理和生成透明、生物意义明确的解释方面表现不足。BioReason通过将DNA基础模型与大型语言模型(LLM)紧密集成,使LLM能够直接解读并推理基因组信息。经过监督微调和强化学习,BioReason学会生成逻辑一致、生物合理的推断。其性能显著提升:基于KEGG的疾病通路预测准确率从86%提高到98%,变异效应预测平均优于强基线15%。该模型能对未见过的生物实体进行推理,并逐步解释决策过程,为可解释、机制性的生物智能提供了新范式。所有数据、代码及检查点均可在https://github.com/bowang-lab/BioReason获取。
原文摘要 · Abstract (English)
Unlocking deep and interpretable biological reasoning from complex genomic data remains a major AI challenge limiting scientific progress. While current DNA foundation models excel at representing sequences, they struggle with multi-step reasoning and lack transparent, biologically meaningful explanations. BioReason addresses this by tightly integrating a DNA foundation model with a large language model (LLM), enabling the LLM to directly interpret and reason over genomic information. Through supervised fine-tuning and reinforcement learning, BioReason learns to produce logical, biologically coherent deductions. It achieves major performance gains, boosting KEGG-based disease pathway prediction accuracy from 86% to 98% and improving variant effect prediction by an average of 15% over strong baselines. BioReason can reason over unseen biological entities and explain its decisions step by step, offering a transformative framework for interpretable, mechanistic AI in biology. All data, code, and checkpoints are available at https://github.com/bowang-lab/BioReason
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。