arXiv:2412.03791cs.LGcs.AI2024-12ICML被引 4

提出直接在坐标空间训练流匹配模型的新方法,跨模态通用且性能更强。

INRFlow: Flow Matching for INRs in Ambient Space

  • 跳过压缩器,直接在环境空间用点级目标训练流匹配模型
  • 在图像、点云、蛋白质结构上均表现优异,超越现有方法
  • 适合需要统一生成模型的跨域研究者使用

流匹配模型已成为图像、视频乃至不规则数据(如3D点云或蛋白质结构)生成建模的强大工具。传统方法通常分两阶段训练:先训练数据压缩器,再在压缩后的隐空间中训练流匹配生成模型。这种两阶段范式限制了跨数据域模型的统一性,因为不同模态需设计特定压缩器架构。为此,我们提出INRFlow,一种在环境空间中直接学习流匹配变换器的领域无关方法。受INRs启发,我们引入条件独立的逐点训练目标,使INRFlow能在坐标空间连续预测。实验证明,INRFlow有效处理图像、3D点云和蛋白质结构等多类数据,在各领域均表现出色,优于可比方法。INRFlow为实现可无缝适配多数据域的领域无关流匹配生成模型迈出了关键一步。

原文摘要 · Abstract (English)

Flow matching models have emerged as a powerful method for generative modeling on domains like images or videos, and even on irregular or unstructured data like 3D point clouds or even protein structures. These models are commonly trained in two stages: first, a data compressor is trained, and in a subsequent training stage a flow matching generative model is trained in the latent space of the data compressor. This two-stage paradigm sets obstacles for unifying models across data domains, as hand-crafted compressors architectures are used for different data modalities. To this end, we introduce INRFlow, a domain-agnostic approach to learn flow matching transformers directly in ambient space. Drawing inspiration from INRs, we introduce a conditionally independent point-wise training objective that enables INRFlow to make predictions continuously in coordinate space. Our empirical results demonstrate that INRFlow effectively handles different data modalities such as images, 3D point clouds and protein structure data, achieving strong performance in different domains and outperforming comparable approaches. INRFlow is a promising step towards domain-agnostic flow matching generative models that can be trivially adopted in different data domains.

流匹配生成模型跨模态INR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。