arXiv:2507.16663cs.CLcs.AI2025-07被引 7

用理解能力提升生成质量,实现多模态模型的自我优化

Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs

  • 利用强理解能力指导弱生成,无须外部数据
  • 生成性能显著提升,理解与生成趋于统一
  • 发现生成与理解协同进化现象,适合模型优化研究者

尽管统一的多模态大模型(MLLMs)旨在融合生成与理解,但普遍存在内部差距:理解能力优于生成能力。大规模评估证实这一非统一性广泛存在,且根源在于生成能力薄弱而非误解。为此,提出基于内部差距的自改进框架,通过更强的理解能力评分生成结果,构建图像数据用于后续微调(如SFT和DPO),有效提升生成性能并促进统一。实验还发现自改进带来协同提升效应:随着生成改善,理解能力更精准识别此前被误判为对齐的错误生成。理论分析表明,生成与理解共享的经验神经正切核(empirical neural tangent kernel)促进学习动态一致,驱动协同进化。据此进一步设计课程学习策略,逐步增强理解与生成,重访预训练模型未充分利用样本,动态扩展微调数据,最终实现性能与统一性的双重提升。

原文摘要 · Abstract (English)

Although unified MLLMs aim to unify generation and understanding, they are considered to exhibit an internal gap, with understanding outperforming generation. Through large-scale evaluation across multiple MLLMs and tasks, we confirm the widespread non-unification of MLLMs, and demonstrate that it indeed stems from weak generation rather than misunderstanding. This finding motivates us to propose a simple yet effective internal gap-based self-improvement framework, which mitigates internal gaps by leveraging stronger understanding to guide weaker generation without relying on any external signals. We validate this strategy through comprehensive experiments: scoring generations with understanding to construct image data for post-training (e.g., SFT and DPO) significantly improves generation while promoting unification. Furthermore, we empirically discover a co-improvement effect of such self-improvement, a phenomenon well known in pre-training but underexplored in post-training. Specifically, as generation improves, understanding becomes more effective at detecting false positives that were previously misclassified as prompt-aligned. To explain this effect, we extend learning dynamic theory to the MLLM setting, showing that the shared empirical neural tangent kernel between generation and understanding encourages aligned learning dynamics, thereby driving co-improvement. This interplay between generation and understanding further motivates a curriculum learning approach for stronger self-improvement: progressively enhanced understanding and generation revisit samples underutilized by pre-trained MLLMs, dynamically expanding post-training data and leading to improved performance and unification.

多模态模型自改进生成统一协同进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。