让生成模型反向提升理解能力,实现视觉认知的自我进化。
Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models

- 将生成作为推理中间步骤,通过自生成视觉思考优化理解
- 在12个基准上显著提升多模态理解性能,生成质量决定感知增益
- 揭示当前模型缺乏稳定任务对齐的自我反思能力,适合研究认知机制者
多模态人工智能的长期目标是构建视觉理解与生成相互增强的统一模型。尽管近期工作如BAGEL、BLIP3已取得显著进展,但实际应用中这种统一仍呈单向性:理解常引导生成,而生成如何支持理解却少有研究。本文重新审视这一不对称性,提出生成到理解(G2U)协同机制,使视觉生成成为显式的中间推理步骤。框架允许模型执行受控生成行为,如细节增强、上下文扩展或结构可视化,生成自洽的视觉思维,并反馈至模型以优化感知,无需重训练或外部工具。在十二个基准上的综合评估表明,这种逆向信息流能持续提升多模态理解能力。我们发现生成保真度限制了感知增益,且不同编辑提示类型影响迁移效率。进一步分析显示,模型虽可生成合理修改,但其自生成视觉思维缺乏稳定的任务对齐,表明当前大型多模态模型尚未具备真正的自我反思能力。该工作揭示统一认知中的缺失机制,提示想象并非理解的终点,而是起点。
原文摘要 · Abstract (English)
The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Despite recent works such as BAGEL, BLIP3o achieves remarkable progress; In practice, however, this unification remains one-directional: understanding routinely guides generation, yet how and why generation can support understanding is rarely investigated. We revisit this asymmetry and propose Generation-to-Understanding (G2U) synergy, where visual generation becomes an explicit intermediate reasoning step. Our framework enables a model to perform controlled generative acts, such as detail enhancement, context expansion or structural visualisation, to produce self-generated visual thoughts, which are then fed back into the model to refine perception without retraining or external tools. Through a comprehensive evaluation on twelve benchmarks, this reversed information flow consistently improves multimodal understanding. We show that generative fidelity bounds perceptual gain and that distinct families of edit prompts govern transfer efficiency. We further analyse whether models can decide what to imagine. While they can produce plausible edits, these self-generated visual thoughts lack stable task alignment, revealing that current large multimodal models fall short of true self-reflection. This work exposes a missing mechanism in unified cognition and suggests that imagination is not the end of understanding but its beginning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。