RNA-GPT用大模型理解RNA序列,帮科研人员快速查资料。
RNA-GPT: Multimodal Generative System for RNA Sequence Understanding
- 将RNA序列编码与大语言模型结合,实现跨模态信息对齐。
- 构建含40万条RNA样本的RNA-QA数据集,支持精准指令微调。
- 适合生物医药研究者,加速RNA功能发现与药物研发。
RNA是承载遗传信息的关键分子,对药物开发和生物技术具有深远影响。然而,海量文献使RNA研究进展受阻。为此,我们提出RNA-GPT,一种多模态RNA对话模型,通过整合RNA序列编码器、线性投影层与先进大语言模型(LLMs),实现精确表征对齐,可处理用户上传的RNA序列并生成简洁准确的回答。基于可扩展训练流程,RNA-GPT利用RNA-QA系统——该系统采用分治策略结合GPT-4o与潜在狄利克雷分配(LDA),从RNACentral自动收集RNA注释,高效生成指令微调样本。实验表明,RNA-GPT能有效应对此类复杂查询,推动RNA研究。此外,我们发布包含407,616条RNA样本的RNA-QA数据集,用于模态对齐与指令调优,进一步提升RNA研究工具的潜力。
原文摘要 · Abstract (English)
RNAs are essential molecules that carry genetic information vital for life, with profound implications for drug development and biotechnology. Despite this importance, RNA research is often hindered by the vast literature available on the topic. To streamline this process, we introduce RNA-GPT, a multi-modal RNA chat model designed to simplify RNA discovery by leveraging extensive RNA literature. RNA-GPT integrates RNA sequence encoders with linear projection layers and state-of-the-art large language models (LLMs) for precise representation alignment, enabling it to process user-uploaded RNA sequences and deliver concise, accurate responses. Built on a scalable training pipeline, RNA-GPT utilizes RNA-QA, an automated system that gathers RNA annotations from RNACentral using a divide-and-conquer approach with GPT-4o and latent Dirichlet allocation (LDA) to efficiently handle large datasets and generate instruction-tuning samples. Our experiments indicate that RNA-GPT effectively addresses complex RNA queries, thereby facilitating RNA research. Additionally, we present RNA-QA, a dataset of 407,616 RNA samples for modality alignment and instruction tuning, further advancing the potential of RNA research tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。