图像生成模型可通用理解视觉信息,表现媲美专业模型。
Image Generators are Generalist Vision Learners

- 将视觉任务转为图像生成,用统一接口实现多任务理解。
- 在2D/3D理解任务上达到当前最优,超越零样本专用模型。
- 轻量微调不损失生成能力,适合构建通用视觉基础模型。
近期研究显示,图像与视频生成模型展现出零样本视觉理解能力,类似大语言模型通过生成预训练获得语言理解与推理的涌现能力。尽管长久以来推测生成视觉内容的能力意味着理解能力,但缺乏充分证据证明生成式视觉模型具备强大理解力。本文表明,图像生成训练扮演了类似LLM预训练的角色,使模型学习到强大且通用的视觉表征,在多种视觉任务中实现顶尖表现。我们提出Vision Banana,通过在原始数据基础上加入少量视觉任务数据,对Nano Banana Pro(NBP)进行指令微调构建通用模型。通过将视觉任务输出空间参数化为RGB图像,将感知问题无缝转化为图像生成任务。Vision Banana在涉及2D和3D理解的多个任务上达到当前最优,超越或匹敌零样本专用模型,包括Segment Anything Model 3的分割任务和Depth Anything系列的度量深度估计。结果仅需轻量级指令微调即可达成,且不损害原模型的图像生成能力。优异表现表明图像生成预训练是通用视觉学习者,同时显示图像生成可作为视觉任务的统一通用接口,类似文本生成在语言理解中的作用。这或标志着计算机视觉的重大范式转变:生成式视觉预训练将在构建生成与理解双功能的基础视觉模型中占据核心地位。
原文摘要 · Abstract (English)
Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited evidence that generative vision models have developed strong understanding capabilities. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable SOTA performance on various vision tasks. We introduce Vision Banana, a generalist model built by instruction-tuning Nano Banana Pro (NBP) on a mixture of its original training data alongside a small amount of vision task data. By parameterizing the output space of vision tasks as RGB images, we seamlessly reframe perception as image generation. Our generalist model, Vision Banana, achieves SOTA results on a variety of vision tasks involving both 2D and 3D understanding, beating or rivaling zero-shot domain-specialists, including Segment Anything Model 3 on segmentation tasks, and the Depth Anything series on metric depth estimation. We show that these results can be achieved with lightweight instruction-tuning without sacrificing the base model's image generation capabilities. The superior results suggest that image generation pretraining is a generalist vision learner. It also shows that image generation serves as a unified and universal interface for vision tasks, similar to text generation's role in language understanding and reasoning. We could be witnessing a major paradigm shift for computer vision, where generative vision pretraining takes a central role in building Foundational Vision Models for both generation and understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。