arXiv:2509.14287q-bio.QMcs.AI2025-09被引 1

让基因序列设计更精准:保持功能属性的几何结构,实现高效生成。

Property-Isometric Variational Autoencoders for Sequence Modeling and Design

  • 构建属性保形的变分自编码框架,用图神经网络和等距正则化保持序列与功能间的几何关系。
  • 在荧光纳米簇和抗菌肽设计中,生成序列使稀有功能类型富集达16.1倍。
  • 适合生物分子设计者、合成生物学研究者,尤其关注复杂多维功能优化的场景。

具有特定功能属性的生物序列设计(如DNA、RNA或肽)在新型纳米材料、生物传感器、抗菌药物等领域有广泛应用。现有模型通常仅依赖二元标签(如结合/非结合),难以优化高维复杂属性,如DNA介导荧光纳米颗粒的目标发射光谱、光化学稳定性及肽对靶标微生物的抗菌活性。为此,我们提出一种几何保形变分自编码框架PrIVAE,学习保留属性空间几何结构的序列潜在表示。具体地,将属性空间建模为高维流形,通过定义合理的距离度量,用最近邻图局部近似其结构,并利用该属性图通过图神经网络编码层和等距正则化引导序列潜在表示。PrIVAE学习到按属性组织的潜在空间,支持通过训练好的解码器理性设计具有期望属性的新序列。我们在两类生成任务上评估框架有效性:(1)设计模板化荧光金属纳米簇的DNA序列;(2)设计抗菌肽。模型在保持高重构准确率的同时,实现了属性有序的潜在空间。除仿真实验外,还基于采样序列开展湿实验设计,在荧光纳米簇中实现稀有属性簇最多16.1倍的富集,验证了框架的实际价值。

原文摘要 · Abstract (English)

Biological sequence design (DNA, RNA, or peptides) with desired functional properties has applications in discovering novel nanomaterials, biosensors, antimicrobial drugs, and beyond. One common challenge is the ability to optimize complex high-dimensional properties such as target emission spectra of DNA-mediated fluorescent nanoparticles, photo and chemical stability, and antimicrobial activity of peptides across target microbes. Existing models rely on simple binary labels (e.g., binding/non-binding) rather than high-dimensional complex properties. To address this gap, we propose a geometry-preserving variational autoencoder framework, called PrIVAE, which learns latent sequence embeddings that respect the geometry of their property space. Specifically, we model the property space as a high-dimensional manifold that can be locally approximated by a nearest neighbor graph, given an appropriately defined distance measure. We employ the property graph to guide the sequence latent representations using (1) graph neural network encoder layers and (2) an isometric regularizer. PrIVAE learns a property-organized latent space that enables rational design of new sequences with desired properties by employing the trained decoder. We evaluate the utility of our framework for two generative tasks: (1) design of DNA sequences that template fluorescent metal nanoclusters and (2) design of antimicrobial peptides. The trained models retain high reconstruction accuracy while organizing the latent space according to properties. Beyond in silico experiments, we also employ sampled sequences for wet lab design of DNA nanoclusters, resulting in up to 16.1-fold enrichment of rare-property nanoclusters compared to their abundance in training data, demonstrating the practical utility of our framework.

序列设计生成模型生物信息学属性保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。