统一模型生成能力未必提升理解,特定任务下才有效
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
- 构建7类30子任务的G2U评估基准,检验生成如何助理解
- 统一模型普遍弱于基础视觉语言模型,生成后回答常降低性能
- 空间推理、多轮思考等任务中生成有增益,体现模型内在偏见
统一多模态模型虽具强大生成能力,但生成是否促进理解尚不明确。现有基准缺乏对生成提升理解的具体任务系统性探索。为此,我们提出UniG2U-Bench,涵盖7种范式和30个子任务,要求不同程度的隐式或显式视觉转换。对30多个模型的广泛评估揭示三大发现:1)统一模型通常弱于基础视觉-语言模型(VLM),生成后回答(GtA)推理普遍劣于直接推理;2)在空间智能、视觉错觉或多轮推理子任务中表现持续提升,表明增强的空间与形状感知及多步中间图像状态有益;3)具有相似推理结构的任务和共享架构的模型表现出相关行为,说明生成-理解耦合诱发了跨任务、预训练数据与模型架构的一致归纳偏见。这些结果凸显需要更丰富的训练数据和新范式以充分释放统一多模态建模潜力。
原文摘要 · Abstract (English)
Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where generation facilitates understanding. To this end, we introduce UniG2U-Bench, a comprehensive benchmark categorizing generation-to-understanding (G2U) evaluation into 7 regimes and 30 subtasks, requiring varying degrees of implicit or explicit visual transformations. Extensive evaluation of over 30 models reveals three core findings: 1) Unified models generally underperform their base Vision-Language Models (VLMs), and Generate-then-Answer (GtA) inference typically degrades performance relative to direct inference. 2) Consistent enhancements emerge in spatial intelligence, visual illusions, or multi-round reasoning subtasks, where enhanced spatial and shape perception, as well as multi-step intermediate image states, prove beneficial. 3) Tasks with similar reasoning structures and models sharing architectures exhibit correlated behaviors, suggesting that generation-understanding coupling induces class-consistent inductive biases over tasks, pretraining data, and model architectures. These findings highlight the necessity for more diverse training data and novel paradigms to fully unlock the potential of unified multimodal modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。