arXiv:2608.22174cs.CV2026-08

探究视觉生成能否提升多模态模型理解能力,发现简单任务有效,复杂任务反而失效。

When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?

论文配图:When Does Visual Generation Help Visual Understanding in Unified Multimodal Models?
图 1 · 摘自论文原文
  • 设计细粒度评估框架VGAU-Diag,分离难度与推理模式影响。
  • 生成辅助在简单任务中提升理解,但复杂任务下可靠性下降。
  • 瓶颈在理解侧而非生成侧,应针对性优化视觉理解能力。

统一多模态模型(UMMs)可同时完成理解与生成任务,核心问题是:视觉生成能否改善理解?现有评估结果矛盾,且受任务难度、推理范式及生成-理解闭环交互的干扰。本文提出VGAU-Diag,一个细粒度评估框架,按难度分层样本,统一评估多种推理范式,并采用Oracle辅助参考协议。分析表明,生成的视觉辅助在简单实例中有帮助,但随着推理复杂度增加,其可靠性下降。Oracle辅助诊断进一步揭示,主要瓶颈通常在视觉理解侧而非生成侧,当前UMMs难以有效利用忠实的视觉辅助。我们还发现,有效的视觉生成应针对理解瓶颈,而非增加更多推理步骤,并识别出从任务无关噪声、到误导性合理引导、再到真正辅助的三阶段转变。这些发现有助于指导更优UMMs的设计。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) can perform both understanding and generation, raising a central question: can visual generation improve understanding? Existing evaluations provide mixed evidence, but confound task difficulty, reasoning paradigms, and the closed-loop interaction between generation and understanding. We introduce VGAU-Diag, a fine-grained evaluation framework for vision generation-assisted understanding. It stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Assisted Reference Protocols. Our analysis shows that generated visual aids help on easier instances but become unreliable as reasoning complexity increases. Oracle-assisted diagnosis further reveals that the main bottleneck often lies on the visual-understanding side rather than the visual-generation side, as current UMMs struggle to leverage even faithful visual aids. We also show that effective visual generation should target visual-understanding bottlenecks rather than add more reasoning steps, and identify a three-stage transition from task-irrelevant noise, to misleading plausible guidance, and finally to useful assistance. These findings would be useful to guide the development of better UMMs.

多模态视觉理解生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。