构建首个跨学科多模态统一评估基准,检验视觉理解与生成的协同能力。
Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark
- 设计双向耦合任务,要求模型用理解指导生成或用生成辅助推理。
- 覆盖8个领域,包含可验证的中间推理步骤和独特真实标签。
- 揭示生成与理解间显著性能差异,适合研究统一多模态模型者参考。
统一多模态模型旨在同时实现视觉理解与生成,但现有评估基准很少考察二者的真实融合。现有评测要么将两种能力分开,要么忽略天然耦合的任务。为弥补这一空白,我们提出Uni-MMMU,一个全面且学科感知的基准,系统展开八个以推理为核心的领域(包括科学、编程、数学和谜题)中生成与理解之间的双向协同。每个任务均双向耦合,要求模型(i)利用概念理解引导精确视觉合成,或(ii)利用生成作为分析推理的认知支架。Uni-MMMU包含可验证的中间推理步骤、唯一真实标签及可复现的文本与视觉输出评分协议。通过对顶尖统一模型、仅生成模型和仅理解模型的广泛评估,我们揭示了显著的性能差距与跨模态依赖关系,提供了关于两者如何相互增强的新见解,并建立了推进统一模型的可靠基础。
原文摘要 · Abstract (English)
Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat the two abilities in isolation or overlook tasks that inherently couple them. To address this gap, we present Uni-MMMU, a comprehensive and discipline-aware benchmark that systematically unfolds the bidirectional synergy between generation and understanding across eight reasoning-centric domains, including science, coding, mathematics, and puzzles. Each task is bidirectionally coupled, demanding models to (i) leverage conceptual understanding to guide precise visual synthesis, or (ii) utilize generation as a cognitive scaffold for analytical reasoning. Uni-MMMU incorporates verifiable intermediate reasoning steps, unique ground truths, and a reproducible scoring protocol for both textual and visual outputs. Through extensive evaluation of state-of-the-art unified, generation-only, and understanding-only models, we reveal substantial performance disparities and cross-modal dependencies, offering new insights into when and how these abilities reinforce one another, and establishing a reliable foundation for advancing unified models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。