提出新方法优化离散流模型,提升序列生成可控性与自然度
Discrete Flow Matching Policy Optimization
- 将采样过程视为多步马尔可夫决策过程,统一强化学习微调框架
- 在基因序列设计任务中,增强增强子活性并保持序列自然性
- 引入总变差正则化防止策略坍塌,适合需高可控性的生成任务
我们提出离散流匹配策略优化(DoMinO),一种适用于广义策略梯度方法的强化学习微调框架,用于离散流匹配(DFM)模型。核心思想是将DFM采样过程视为多步马尔可夫决策过程,从而将奖励最大化重构为稳健的强化学习目标。该方法不仅保留原始采样器,还避免了以往方法中使用的有偏辅助估计和似然代理。为防止策略坍塌,引入新的总变差正则化项,使微调分布贴近预训练分布。理论上,我们建立了DoMinO的离散化误差上界及正则项的可处理上界。实验在调控DNA序列设计任务上验证,DoMinO在预测增强子活性和序列自然度上优于现有最佳奖励驱动基线,正则化进一步提升了与自然序列分布的一致性,同时保持强功能性能。结果表明,DoMinO是可控离散序列生成的有效框架。
原文摘要 · Abstract (English)
We introduce Discrete flow Matching policy Optimization (DoMinO), a unified framework for Reinforcement Learning (RL) fine-tuning Discrete Flow Matching (DFM) models under a broad class of policy gradient methods. Our key idea is to view the DFM sampling procedure as a multi-step Markov Decision Process. This perspective provides a simple and transparent reformulation of fine-tuning reward maximization as a robust RL objective. Consequently, it not only preserves the original DFM samplers but also avoids biased auxiliary estimators and likelihood surrogates used by many prior RL fine-tuning methods. To prevent policy collapse, we also introduce new total-variation regularizers to keep the fine-tuned distribution close to the pretrained one. Theoretically, we establish an upper bound on the discretization error of DoMinO and tractable upper bounds for the regularizers. Experimentally, we evaluate DoMinO on regulatory DNA sequence design. DoMinO achieves stronger predicted enhancer activity and better sequence naturalness than the previous best reward-driven baselines. The regularization further improves alignment with the natural sequence distribution while preserving strong functional performance. These results establish DoMinO as an useful framework for controllable discrete sequence generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。