通过拉回几何实现数据流形上的生成建模,支持高效采样与精准插值。
Pullback Flow Matching on Data Manifolds
- 利用拉回几何与等距学习,保持数据流形结构的同时设计可调控的隐空间。
- 在蛋白质动力学与序列数据上生成具有特定性质的新蛋白,提升生成质量。
- 适用于药物发现与材料科学,适合需要定向生成新样本的研究场景。
我们提出拉回流匹配(Pullback Flow Matching, PFM),一种面向数据流形的生成建模新框架。不同于现有方法需假设或学习受限的闭式流形映射来训练黎曼流匹配(RFM)模型,PFM 利用拉回几何与等距学习,在保持底层流形几何结构的同时,实现高效的生成与精确的隐空间插值。该方法不仅支持在数据流形上实现闭式映射,还允许通过在数据与隐空间上设定度量来设计灵活的隐空间。通过引入神经微分方程增强等距学习,并提出可扩展的训练目标,显著优化了隐空间的插值性能,从而提升流形学习与生成效果。我们在合成数据、蛋白质动力学及蛋白质序列数据上验证了 PFM 的有效性,成功生成具有特定性质的新蛋白,展现出在药物发现与材料科学中生成具备特定功能新样本的强大潜力。
原文摘要 · Abstract (English)
We propose Pullback Flow Matching (PFM), a novel framework for generative modeling on data manifolds. Unlike existing methods that assume or learn restrictive closed-form manifold mappings for training Riemannian Flow Matching (RFM) models, PFM leverages pullback geometry and isometric learning to preserve the underlying manifold's geometry while enabling efficient generation and precise interpolation in latent space. This approach not only facilitates closed-form mappings on the data manifold but also allows for designable latent spaces, using assumed metrics on both data and latent manifolds. By enhancing isometric learning through Neural ODEs and proposing a scalable training objective, we achieve a latent space more suitable for interpolation, leading to improved manifold learning and generative performance. We demonstrate PFM's effectiveness through applications in synthetic data, protein dynamics and protein sequence data, generating novel proteins with specific properties. This method shows strong potential for drug discovery and materials science, where generating novel samples with specific properties is of great interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。