通过保守搜索提升生物序列设计的强化学习鲁棒性
Improved Off-policy Reinforcement Learning in Biological Sequence Design
- 基于噪声注入与去噪机制,限制策略在可信区域探索
- 在DNA、RNA、蛋白质等任务中均超越现有方法发现高分序列
- 动态调整保守程度,适配代理模型的不确定性
由于搜索空间巨大且评估预算有限,设计具有特定功能的生物序列极具挑战。尽管强化学习利用代理模型可快速评估奖励,但训练数据不足会导致代理模型在分布外输入上出现误设。为此,我们提出一种新型离线策略搜索方法——δ-保守搜索,通过将高分离线序列进行随机掩码(概率为δ)后由策略去噪,限制探索范围至可靠区域,从而增强鲁棒性。同时,根据每个数据点的代理模型不确定性动态调整δ值,使保守程度与模型置信度对齐。实验表明,该方法在多种任务(包括DNA、RNA、蛋白和肽段设计)中持续提升离线策略训练效果,显著优于现有机器学习方法,能更有效地发现高分序列。
原文摘要 · Abstract (English)
Designing biological sequences with desired properties is challenging due to vast search spaces and limited evaluation budgets. Although reinforcement learning methods use proxy models for rapid reward evaluation, insufficient training data can cause proxy misspecification on out-of-distribution inputs. To address this, we propose a novel off-policy search, $δ$-Conservative Search, that enhances robustness by restricting policy exploration to reliable regions. Starting from high-score offline sequences, we inject noise by randomly masking tokens with probability $δ$, then denoise them using our policy. We further adapt $δ$ based on proxy uncertainty on each data point, aligning the level of conservativeness with model confidence. Experimental results show that our conservative search consistently enhances the off-policy training, outperforming existing machine learning methods in discovering high-score sequences across diverse tasks, including DNA, RNA, protein, and peptide design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。