通过结构化耦合与潜变量编辑,实现生物序列的灵活高效生成。
Flexible Flows for Biological Sequence Design

- 设计结构化耦合机制,引导生成更符合生物学特性的序列。
- 支持可变长度序列生成,且在多个任务上达到顶尖性能。
- 适合需要精细控制的生物序列设计,如药物肽开发。
生物序列设计需在庞大离散空间中满足严格的进化与生物物理约束。离散流匹配(DFM)提供了一种生成框架,但现有方法依赖无生物学意义的耦合,难以灵活处理可变长度序列和细粒度控制。本文提出一种编码序列元素领域偏好信息的结构化耦合,无需修改流目标或训练流程即可将源分布偏向合理区域。在此基础上,引入基于潜变量编辑的速率参数化,通过共享全局潜变量建模可变长度生成,类似潜变量模型但保持可计算性。进一步提出无分类器潜变量引导机制,在连续潜空间中实现一致生成方向,并采用狄利克雷先验温度缩放,实现测试时对编辑操作的可控调节。方法在多种生物序列任务中表现卓越,涵盖密度估计、无条件与条件DNA生成、肽序列生成等。
原文摘要 · Abstract (English)
Designing functional biological sequences requires navigating vast discrete spaces under strict evolutionary and biophysical constraints. Discrete Flow Matching (DFM) offers a generative framework over such spaces, but existing approaches rely on biologically uninformative couplings and offer limited flexibility for variable-length sequence generation and fine-grained control. We propose a structured coupling that encodes domain-specific preferences among sequence elements, biasing the source distribution toward plausible regions without modifying the flow objective or training procedure. Building on this, we introduce a latent edit-based rate parameterization that models variable-length generation via edit operations conditioned on a shared global latent, akin to a latent variable model, while remaining tractable. We further introduce a latent classifier-free guidance mechanism that steers generation coherently in continuous latent space, along with Dirichlet-prior temperature scaling for test-time control over edit operations. Our method achieves state-of-the-art performance across diverse biological sequence tasks, including density estimation, unconditional and conditional DNA sequence generation, and peptide sequence generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。