arXiv:2410.13848cs.CVcs.AI2024-10CVPR被引 472

Janus通过分离视觉编码路径,提升多模态理解与生成统一模型的表现。

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

论文配图:Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
图 1 · 摘自论文原文
  • 将视觉编码拆分为独立路径,避免理解与生成任务冲突。
  • 在多个基准上超越现有统一模型,接近专用模型性能。
  • 适合构建下一代灵活高效的多模态通用模型。

本文提出Janus,一种统一多模态理解与生成的自回归框架。以往研究通常使用单一视觉编码器处理两类任务(如Chameleon),但因理解与生成对信息粒度要求不同,导致性能受限,尤其在多模态理解方面。为此,Janus将视觉编码解耦为独立路径,仍采用统一Transformer架构进行处理。该设计缓解了视觉编码器在两类任务间的角色冲突,同时增强框架灵活性,使理解与生成模块可自主选择最优编码方式。实验表明,Janus超越现有统一模型,在多数任务上达到或超过专用模型性能。其简洁性、高灵活性和有效性使其成为下一代统一多模态模型的有力候选。

原文摘要 · Abstract (English)

In this paper, we introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder's roles in understanding and generation, but also enhances the framework's flexibility. For instance, both the multimodal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models.

多模态统一模型视觉编码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。