无需已知序列即可设计蛋白结合RNA,突破传统依赖瓶颈。
BAnG: Bidirectional Anchored Generation for Conditional RNA Design
- 双向锚定生成法聚焦功能基序与上下文关系
- 在合成与真实生物数据上均实现高效条件生成
- 适合缺乏先验数据的蛋白-核酸设计任务
设计能与特定蛋白质相互作用的RNA分子是实验与计算生物学中的关键挑战。现有计算方法通常需大量已知的相互作用RNA序列或详细的RNA结构信息,限制了实际应用。为解决这一问题,我们提出基于深度学习的RNA-BAnG模型,可在无需这些前提条件下生成蛋白结合型RNA序列。核心是新型生成方法——双向锚定生成(BAnG),其基于蛋白结合RNA常包含嵌入在广泛序列背景中的功能结合基序这一观察。我们首先在含类似局部基序的合成任务中验证该方法,证明其优于现有生成方法;随后在真实生物序列上评估,证实其在给定结合蛋白时进行条件性RNA序列设计的有效性。
原文摘要 · Abstract (English)
Designing RNA molecules that interact with specific proteins is a critical challenge in experimental and computational biology. Existing computational approaches require a substantial amount of previously known interacting RNA sequences for each specific protein or a detailed knowledge of RNA structure, restricting their utility in practice. To address this limitation, we develop RNA-BAnG, a deep learning-based model designed to generate RNA sequences for protein interactions without these requirements. Central to our approach is a novel generative method, Bidirectional Anchored Generation (BAnG), which leverages the observation that protein-binding RNA sequences often contain functional binding motifs embedded within broader sequence contexts. We first validate our method on generic synthetic tasks involving similar localized motifs to those appearing in RNAs, demonstrating its benefits over existing generative approaches. We then evaluate our model on biological sequences, showing its effectiveness for conditional RNA sequence design given a binding protein.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。