arXiv:2509.24897cs.AI2025-09中稿 · CVPR被引 31

测试统一模型能否真正实现理解与生成的相互促进。

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

  • 设计双向协同评估框架,分步检验理解如何指导生成、生成如何辅助理解。
  • 12个主流统一模型在关键任务上表现平平,协同效应不明显。
  • 适合关注多模态模型内在机制与训练方法的研究者参考。

将视觉理解与生成整合到统一的多模态模型中是迈向通用人工智能的重要一步。然而,现有基准无法回答一个根本问题:这种架构统一是否真能促成两种能力之间的协同作用?当前评估方式主要孤立地测试理解与生成能力,难以判断统一模型能否利用理解能力提升生成质量,或通过生成模拟深化理解。为此,我们提出RealUnify,一个专为评估双向能力协同而设计的基准。该基准包含1,000个经人工标注的实例,覆盖10个类别和32个子任务,围绕两大核心维度:1)理解增强生成(需常识或逻辑推理引导图像生成);2)生成增强理解(需心理模拟或重构被变换/打乱的视觉输入来完成推理任务)。其关键贡献在于双评估协议——结合端到端直接评估与分阶段诊断式评估,可精准识别性能瓶颈是源于核心能力不足,还是融合失败。对12个领先统一模型和6个专用基线的规模化评估表明,当前统一模型仍难以实现有效协同,说明仅靠架构统一并不足够。这凸显了需要新的训练策略与归纳偏置,才能充分释放统一建模的潜力。

原文摘要 · Abstract (English)

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question remains unanswered by existing benchmarks: does this architectural unification actually enable synergetic interaction between the constituent capabilities? Existing evaluation paradigms, which primarily assess understanding and generation in isolation, are insufficient for determining whether a unified model can leverage its understanding to enhance its generation, or use generative simulation to facilitate deeper comprehension. To address this critical gap, we introduce RealUnify, a benchmark specifically designed to evaluate bidirectional capability synergy. RealUnify comprises 1,000 meticulously human-annotated instances spanning 10 categories and 32 subtasks. It is structured around two core axes: 1) Understanding Enhances Generation, which requires reasoning (e.g., commonsense, logic) to guide image generation, and 2) Generation Enhances Understanding, which necessitates mental simulation or reconstruction (e.g., of transformed or disordered visual inputs) to solve reasoning tasks. A key contribution is our dual-evaluation protocol, which combines direct end-to-end assessment with a diagnostic stepwise evaluation that decomposes tasks into distinct understanding and generation phases. This protocol allows us to precisely discern whether performance bottlenecks stem from deficiencies in core abilities or from a failure to integrate them. Through large-scale evaluations of 12 leading unified models and 6 specialized baselines, we find that current unified models still struggle to achieve effective synergy, indicating that architectural unification alone is insufficient. These results highlight the need for new training strategies and inductive biases to fully unlock the potential of unified modeling.

多模态协同评估统一模型生成理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。