arXiv:2602.02140cs.CL2026-02被引 10

测试统一多模态模型的理解与生成能力是否真正协同

Quantifying the Gap between Understanding and Generation within Unified Multimodal Models

  • 设计双向评估基准GapEval,量化理解与生成的差距
  • 发现多数模型在两种能力间存在持续鸿沟
  • 揭示知识在不同模态间不连贯,适合研究模型认知机制

统一多模态模型(UMM)在理解和生成任务上均取得显著进展。然而,这两种能力是否在单一模型中真正协同仍不明确。为此,我们提出GapEval——一个双向基准,用于量化理解与生成能力之间的差距,并衡量两种“统一”方向的认知一致性。每个问题均可通过图像和文本两种模态回答,实现对模型双向推理能力和跨模态一致性的对称评估。实验表明,无论架构如何,众多UMM在两个方向间均存在持久差距,说明当前模型仅实现表层统一,而非深层认知融合。为进一步探究机制,我们从知识操作角度开展实证研究,结果表明,模型中的知识往往彼此割裂,模态间的知识涌现与能力发展不同步,为后续探索提供方向。

原文摘要 · Abstract (English)

Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are genuinely aligned and integrated within a single model remains unclear. To investigate this question, we introduce GapEval, a bidirectional benchmark designed to quantify the gap between understanding and generation capabilities, and quantitatively measure the cognitive coherence of the two "unified" directions. Each question can be answered in both modalities (image and text), enabling a symmetric evaluation of a model's bidirectional inference capability and cross-modal consistency. Experiments reveal a persistent gap between the two directions across a wide range of UMMs with different architectures, suggesting that current models achieve only surface-level unification rather than deep cognitive convergence of the two. To further explore the underlying mechanism, we conduct an empirical study from the perspective of knowledge manipulation to illustrate the underlying limitations. Our findings indicate that knowledge within UMMs often remains disjoint. The capability emergence and knowledge across modalities are unsynchronized, paving the way for further exploration.

多模态认知评估模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。