arXiv:2609.07884cs.CV2026-09

测试图像生成器能否零样本理解视觉任务,发现它在推理上更鲁棒但不如专用模型高效。

Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation

论文配图:Are Image Generators Zero-Shot Perceivers? A Rigorous Evaluation
图 1 · 摘自论文原文
  • 将深度估计等任务转为文本提示的生成任务进行评估
  • 20个模型在11个基准上测试,生成模型在分布外数据表现更好
  • 适合关注跨模态推理与生成理解的研究者

近期研究显示,轻量级指令微调可使图像生成器在多个视觉感知任务中达到顶尖水平。受此启发,我们探究图像生成器在零样本设置下于公开视觉感知基准上的能力极限。本文提出ProbeGen基准,将单目深度估计、指代分割和目标计数等任务转化为由文本提示驱动的条件生成任务,并对20种模型(包括专有及开源图像生成器、专用感知模型、多模态大模型)在11个已发表基准上进行比较。结果表明,预训练图像生成器具备可观测的零样本感知能力,但存在明显权衡:专用模型在分布内任务中准确率和效率更高,而生成模型在分布外场景下更具鲁棒性,且在组合语义推理方面表现更优。本研究旨在推动零样本生成式感知成为有意义的研究方向,并为生成与理解交叉领域提供基础支持。

原文摘要 · Abstract (English)

Recent work, such as Vision Banana, shows that lightweight instruction tuning can enable an image generator to achieve state-of-the-art performance across multiple visual perception tasks. Motivated by this perspective, we ask how far image generators can go on public visual perception benchmarks in a zero-shot setting. We introduce ProbeGen, a benchmark for zero-shot generative perception that casts monocular depth estimation, referring/reasoning segmentation, and object counting as conditional generation tasks specified through text prompts, and compares 20 models in total---including proprietary and open-weight image generators, specialist perception models, and MLLMs---across 11 published benchmarks. We observe that pretrained image generators show measurable zero-shot perceptual competence, but with a clear trade-off: specialist models remain stronger for in-distribution accuracy and efficiency, while generative models are often more robust under distribution shift and better at compositional semantic reasoning. We hope this study helps establish zero-shot generative perception as a meaningful research direction and provides a useful foundation for future work at the intersection of visual generation and understanding.

图像生成零样本视觉理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。