用强化学习设计高功能基因调控序列,兼顾生物规律与多样性。
Regulatory DNA sequence Design with Reinforcement Learning
- 用强化学习微调预训练生成模型,结合转录因子结合位点增减策略。
- 在酵母和人类细胞中成功生成高表达活性的启动子与增强子。
- 融合生物先验知识,避免局部最优,适合基因工程与治疗应用。
顺式调控元件(CREs)如启动子和增强子是调控基因表达的短DNA序列,其功能依赖于特定核苷酸序列,尤其是转录因子结合位点(TFBS)。现有CRE设计方法存在两大缺陷:(1)依赖迭代优化,易陷入局部最优;(2)缺乏生物学先验指导。本文提出一种生成式方法,利用强化学习(RL)微调预训练自回归(AR)模型,通过计算推断奖励机制模拟激活型TFBS添加与抑制型TFBS移除,融入RL过程。我们在两种酵母培养条件下的启动子设计及三种人类细胞类型的增强子设计任务上评估该方法,结果表明其能生成高功能性的CREs并保持序列多样性。代码已开源:https://github.com/yangzhao1230/TACO。
原文摘要 · Abstract (English)
Cis-regulatory elements (CREs), such as promoters and enhancers, are relatively short DNA sequences that directly regulate gene expression. The fitness of CREs, measured by their ability to modulate gene expression, highly depends on the nucleotide sequences, especially specific motifs known as transcription factor binding sites (TFBSs). Designing high-fitness CREs is crucial for therapeutic and bioengineering applications. Current CRE design methods are limited by two major drawbacks: (1) they typically rely on iterative optimization strategies that modify existing sequences and are prone to local optima, and (2) they lack the guidance of biological prior knowledge in sequence optimization. In this paper, we address these limitations by proposing a generative approach that leverages reinforcement learning (RL) to fine-tune a pre-trained autoregressive (AR) model. Our method incorporates data-driven biological priors by deriving computational inference-based rewards that simulate the addition of activator TFBSs and removal of repressor TFBSs, which are then integrated into the RL process. We evaluate our method on promoter design tasks in two yeast media conditions and enhancer design tasks for three human cell types, demonstrating its ability to generate high-fitness CREs while maintaining sequence diversity. The code is available at https://github.com/yangzhao1230/TACO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。