arXiv:2411.07975cs.CVcs.AI2024-11CVPR被引 160

一个模型搞定图文理解与生成,性能超越现有统一方案。

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

  • 将自回归语言模型与修正流结合,架构极简无需复杂改造。
  • 图文任务表现媲美甚至超过专用模型,统一模型更优。
  • 适合追求高效多模态系统的研发人员使用。

我们提出JanusFlow,一种统一图像理解与生成的单模型框架。该框架采用极简架构,将自回归语言模型与修正流(rectified flow)融合,后者是生成建模的前沿方法。关键发现表明,修正流可直接在大型语言模型框架中训练,无需复杂的结构改动。为提升统一模型性能,我们采用两项策略:(i) 解耦理解与生成编码器,(ii) 在联合训练中对齐两者表征。大量实验表明,JanusFlow在标准基准上达到或超越专用模型的表现,同时显著优于现有统一方法。本工作推动了更高效、更通用的视觉-语言模型发展。

原文摘要 · Abstract (English)

We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches across standard benchmarks. This work represents a step toward more efficient and versatile vision-language models.

多模态生成模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。