arXiv:2410.05292cs.LGcs.AI2024-10被引 3

用大模型直接生成连续数据,提升生成效率和上下文感知能力。

CaLMFlow: Volterra Flow Matching using Causal Language Models

  • 将流匹配建模为伏尔泰拉积分方程,通过分时空分词实现序列建模
  • 在单细胞扰动预测等任务上优于传统基于ODE的方法,支持文本上下文输入
  • 适合需要高维数据生成与上下文理解的科研场景,如生物医学建模

我们提出CaLMFlow(因果语言模型用于流匹配),将流匹配表述为伏尔泰拉积分方程(VIE),利用大语言模型(LLM)实现连续数据生成。该方法通过时空分词,将流匹配转化为序列建模任务,打通离散语言建模与连续生成建模的鸿沟。其在高维数据处理中表现高效,优于依赖微分方程求解器的条件流匹配(CFM)。我们在合成数据和真实世界数据上验证了有效性,包括单细胞扰动响应预测,展示了其对文本上下文的融合能力及对未见条件的泛化性能。结果表明,以大模型驱动的流匹配是生成建模中具有前景的新范式,具备更好的可扩展性、灵活性和上下文感知能力。

原文摘要 · Abstract (English)

We introduce CaLMFlow (Causal Language Models for Flow Matching), a novel framework that casts flow matching as a Volterra integral equation (VIE), leveraging the power of large language models (LLMs) for continuous data generation. CaLMFlow enables the direct application of LLMs to learn complex flows by formulating flow matching as a sequence modeling task, bridging discrete language modeling and continuous generative modeling. Our method implements tokenization across space and time, thereby solving a VIE over these domains. This approach enables efficient handling of high-dimensional data and outperforms ODE solver-dependent methods like conditional flow matching (CFM). We demonstrate CaLMFlow's effectiveness on synthetic and real-world data, including single-cell perturbation response prediction, showcasing its ability to incorporate textual context and generalize to unseen conditions. Our results highlight LLM-driven flow matching as a promising paradigm in generative modeling, offering improved scalability, flexibility, and context-awareness.

生成模型大模型流匹配生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。