arXiv:2506.10395cs.CVcs.AI2025-06被引 6

Pisces用分离编码器统一图像理解与生成,性能超越专用模型。

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

  • 采用解耦视觉编码架构,分别优化理解与生成所需特征。
  • 在20多个图像理解基准上表现优异,GenEval生成评测达领先水平。
  • 适合研究统一多模态模型的学者,尤其关注生成与理解协同的场景。

近年来,大语言模型的发展推动了多模态基础模型在统一框架下同时处理图像理解与生成的能力。尽管取得进展,统一模型在任一任务上的表现仍常落后于专用模型。其核心挑战在于图像理解与生成所需的视觉特征本质不同,且训练过程差异显著。本文提出Pisces,一种自回归多模态基础模型,通过创新的解耦视觉编码架构和针对多模态生成优化的训练策略解决此问题。结合精细的数据筛选、预训练与微调流程,Pisces在图像理解与生成两方面均实现竞争力表现。我们在超过20个公开图像理解基准上评估该模型,展现出广泛任务下的强性能;在广泛使用的图像生成基准GenEval上,Pisces亦展现稳健生成能力。大量分析揭示了图像理解与生成间的协同关系,以及使用独立视觉编码器的优势,推动了统一多模态模型的发展。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite these gains, unified models often underperform compared to specialized models in either task. A key challenge in developing unified models lies in the inherent differences between the visual features needed for image understanding versus generation, as well as the distinct training processes required for each modality. In this work, we introduce Pisces, an auto-regressive multimodal foundation model that addresses this challenge through a novel decoupled visual encoding architecture and tailored training techniques optimized for multimodal generation. Combined with meticulous data curation, pretraining, and finetuning, Pisces achieves competitive performance in both image understanding and image generation. We evaluate Pisces on over 20 public benchmarks for image understanding, where it demonstrates strong performance across a wide range of tasks. Additionally, on GenEval, a widely adopted benchmark for image generation, Pisces exhibits robust generative capabilities. Our extensive analysis reveals the synergistic relationship between image understanding and generation, and the benefits of using separate visual encoders, advancing the field of unified multimodal models.

多模态图像生成自回归统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。