对比13种RNA语言模型,发现结构与功能预测难以兼顾。
A Comparative Review of RNA Language Models
- 按训练数据分为三类,对比13种RNA模型性能
- 结构预测强的模型功能分类表现差,反之亦然
- 提示需更平衡的无监督训练,适合生物信息研究者
鉴于蛋白质语言模型在结构与功能推断中的有效性,近年来对RNA语言模型的关注日益增加。然而,这些模型常缺乏统一的评估标准。本文将RNA语言模型分为三类:基于多种RNA类型(特别是非编码RNA)预训练、针对特定功能RNA设计、以及统一处理RNA与DNA或蛋白质的模型。我们对比了13种RNA LMs,以及3种DNA和1种蛋白质模型,在零样本条件下预测RNA二级结构和功能分类的表现。结果表明,结构预测性能优异的模型在功能分类任务中往往表现较差,反之亦然,提示需要更均衡的无监督训练策略。
原文摘要 · Abstract (English)
Given usefulness of protein language models (LMs) in structure and functional inference, RNA LMs have received increased attentions in the last few years. However, these RNA models are often not compared against the same standard. Here, we divided RNA LMs into three classes (pretrained on multiple RNA types (especially noncoding RNAs), specific-purpose RNAs, and LMs that unify RNA with DNA or proteins or both) and compared 13 RNA LMs along with 3 DNA and 1 protein LMs as controls in zero-shot prediction of RNA secondary structure and functional classification. Results shows that the models doing well on secondary structure prediction often perform worse in function classification or vice versa, suggesting that more balanced unsupervised training is needed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。