arXiv:2410.16212cs.AIcs.LG2024-10被引 28

对比多个RNA大模型在二级结构预测中的表现,发现两个模型更优。

Comprehensive benchmarking of large language models for RNA secondary structure prediction

  • 在统一框架下评估多个预训练RNA大模型的性能
  • 两个模型在基准数据集上显著优于其他模型
  • 低相似性场景下泛化能力仍是关键挑战

受大语言模型在DNA和蛋白质领域成功启发,近期涌现出多个针对RNA的LLM。RNA-LLM利用大规模RNA序列数据,以自监督方式学习每个核苷酸的语义丰富向量表示,旨在提升数据密集型下游任务的性能。其中,二级结构预测是揭示RNA功能机制的基础任务。本文对多个预训练RNA-LLM进行了全面实验分析,在统一深度学习框架下比较其在二级结构预测任务中的表现。通过在不同泛化难度的基准数据集上测试,结果表明有两个模型明显优于其他模型,并揭示了在低同源性场景下泛化能力仍面临显著挑战。

原文摘要 · Abstract (English)

Inspired by the success of large language models (LLM) for DNA and proteins, several LLM for RNA have been developed recently. RNA-LLM uses large datasets of RNA sequences to learn, in a self-supervised way, how to represent each RNA base with a semantically rich numerical vector. This is done under the hypothesis that obtaining high-quality RNA representations can enhance data-costly downstream tasks. Among them, predicting the secondary structure is a fundamental task for uncovering RNA functional mechanisms. In this work we present a comprehensive experimental analysis of several pre-trained RNA-LLM, comparing them for the RNA secondary structure prediction task in an unified deep learning framework. The RNA-LLM were assessed with increasing generalization difficulty on benchmark datasets. Results showed that two LLM clearly outperform the other models, and revealed significant challenges for generalization in low-homology scenarios.

RNA结构大模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。