arXiv:2503.09085q-bio.BMcs.LG2025-03被引 5

用可微折叠优化RNA结构模型参数,性能提升23个数量级。

Differentiable Folding for Nearest Neighbor Model Optimization

  • 基于可微折叠技术,直接计算参数梯度实现高效优化。
  • 新参数集使真实序列-结构对预测概率提升超23个数量级。
  • 适合需要高精度RNA模型的研究者,支持新数据灵活接入。

近邻邻居模型是RNA二级结构形成最主流的热力学模型,也是RNA结构预测与序列设计的基础。当前的函数形式(Turner 2004)包含约13,000个热力学参数,将其拟合到实验和结构数据上计算成本极高。本文利用近期发展的可微折叠技术——一种直接计算RNA折叠算法梯度的方法,提出一种高效、可扩展且灵活的参数优化方法,结合已知的RNA结构与热力学实验数据。该方法获得的参数集在所有指标上均显著优于现有基线,单个RNA家族的真实序列-结构对平均预测概率提升超过23个数量级。该框架为构建更优的RNA模型提供了路径,支持新实验数据的灵活引入、新损失函数定义、大规模训练集使用,甚至可作为模块嵌入更大深度学习流程中。我们还发布了新数据库RNAometer,包含小分子RNA模型系统的实验稳定性数据。

原文摘要 · Abstract (English)

The Nearest Neighbor model is the $\textit{de facto}$ thermodynamic model of RNA secondary structure formation and is a cornerstone of RNA structure prediction and sequence design. The current functional form (Turner 2004) contains $\approx13,000$ underlying thermodynamic parameters, and fitting these to both experimental and structural data is computationally challenging. Here, we leverage recent advances in $\textit{differentiable folding}$, a method for directly computing gradients of the RNA folding algorithms, to devise an efficient, scalable, and flexible means of parameter optimization that uses known RNA structures and thermodynamic experiments. Our method yields a significantly improved parameter set that outperforms existing baselines on all metrics, including an increase in the average predicted probability of ground-truth sequence-structure pairs for a single RNA family by over 23 orders of magnitude. Our framework provides a path towards drastically improved RNA models, enabling the flexible incorporation of new experimental data, definition of novel loss terms, large training sets, and even treatment as a module in larger deep learning pipelines. We make available a new database, RNAometer, with experimentally-determined stabilities for small RNA model systems.

RNA建模可微折叠参数优化生物计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。