OmicsLM让大模型直接理解基因表达数据并用自然语言解释生物意义。
OmicsLM: A Multimodal Large Language Model for Multi-Sample Omics Reasoning

- 将基因表达数据转为连续向量接入大模型,支持多样本和自然语言指令联合处理。
- 在真实数据上表现优于专用模型和通用大模型,尤其擅长语言引导的生物推理。
- 适用于生物研究者、临床分析人员,助力从表达数据中挖掘深层生物学见解。
解析转录组数据是现代生物学中最常见的分析任务之一。然而,现有模型要么仅处理表达谱但不生成自然语言解释,要么仅在语言层面推理却缺乏对定量组学数据的直接访问。我们提出OmicsLM,一个连接定量组学谱与自然语言生物任务的多模态大模型。该模型将每个转录组谱表示为大模型上下文中的紧凑连续表征,既保留定量表达信号,又支持自然语言指令、显式基因提及及多个交错的生物样本在单一上下文中处理。我们在超过550万条指令跟随样本上训练,涵盖70多种任务类型,融合连续转录组输入、通过多样化语言模板呈现的实验数据,以及自由文本的生物知识与问答数据。任务包括细胞类型注释、扰动预测、临床预测、通路推理和开放性生物问题问答。现有基准仅评估谱级预测或纯文本生物问答,未覆盖语言引导的多样本真实表达谱推理。为此,我们构建了基于真实GEO研究的GEO-OmicsQA基准。结果表明,OmicsLM可直接使用表达谱,在谱级任务上表现媲美专用组学模型,同时在语言引导的生物推理任务上显著优于组学专用模型和通用大模型。
原文摘要 · Abstract (English)
Interpreting transcriptomic data is one of the most common analytical tasks in modern biology. Yet most current models either consume expression profiles without producing natural-language biological explanations, or reason in language without direct access to quantitative omics measurements. We introduce OmicsLM, a multimodal LLM that connects quantitative omics profiles with natural-language biological tasks. OmicsLM represents each transcriptomic profile as a compact continuous representation within the LLM context. This interface preserves quantitative expression signal while allowing natural-language instructions, explicit gene mentions, and multiple interleaved biological samples to be processed together in one model context. We train OmicsLM on more than 5.5 million instruction-following examples spanning over 70 task types, combining continuous transcriptomic inputs, experimental data rendered through diverse language templates, and free-text biological knowledge and question-answering data. This mixture covers cell type annotation, perturbation prediction, clinical prediction, pathway reasoning, and open-ended biological question answering. Existing benchmarks evaluate either profile-level prediction or text-only biological QA, leaving language-guided, multi-sample reasoning over real expression profiles unmeasured. To close this gap, we introduce GEO-OmicsQA, a benchmark for multi-sample biological question answering built from real Gene Expression Omnibus (GEO) studies. We demonstrate that OmicsLM can use expression profiles directly and perform comparably to specialized omics models on profile-level tasks, while outperforming both omics-specialized models and general LLMs on language-guided biological reasoning over expression data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。