评测多模态模型在推理驱动图像生成中的理解与生成一致性
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- 设计三类任务评估统一模型的理解-生成一致性
- 发现统一模型在推理任务中仍存在理解与生成脱节
- 适合研究多模态推理与视觉生成对齐的学者参考
统一多模态模型将大语言模型的推理能力与图像理解及生成结合,展现出先进多模态智能的巨大潜力。然而,社区仍缺乏严谨的以推理为核心的基准测试,难以系统评估理解与生成之间的对齐程度及其在复杂视觉任务中的泛化能力。为此,我们提出GIR-Bench,一个涵盖三个互补视角的综合性基准。首先,评估理解-生成一致性(GIR-Bench-UGC),检验模型在理解和生成任务中是否一致运用相同知识;其次,评估需逻辑约束与隐含知识的推理驱动文本到图像生成(GIR-Bench-T2I);第三,评估多步推理在编辑任务中的表现(GIR-Bench-Edit)。每个子集均设计特定评估流程,实现细粒度、可解释的评估,并缓解当前普遍存在的多模态大模型自评(MLLM-as-a-Judge)带来的偏差。对多种统一模型和仅生成系统的大规模消融实验表明:尽管统一模型更擅长推理驱动的视觉任务,但其理解与生成之间仍存在持续性差距。GIR-Bench的数据与代码已公开于https://github.com/HKUST-LongGroup/GIR-Bench。
原文摘要 · Abstract (English)
Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the alignment between understanding and generation, and their generalization potential in complex visual tasks. To this end, we introduce GIR-Bench, a comprehensive benchmark that evaluates unified models across three complementary perspectives. Firstly, we investigate understanding-generation consistency (GIR-Bench-UGC), asking whether models can consistently leverage the same knowledge in both understanding and generation tasks. Secondly, we investigate whether models can perform reasoning-centric text-to-image generation that requires applying logical constraints and implicit knowledge to generate faithful visual content (GIR-Bench-T2I). Thirdly, we evaluate whether models can handle multi-step reasoning in editing (GIR-Bench-Edit). For each subset, we carefully design different task-specific evaluation pipelines tailored for each task. This enables fine-grained and interpretable evaluation while mitigating biases from the prevalent MLLM-as-a-Judge paradigm. Extensive ablations over various unified models and generation-only systems have shown that: Although unified models are more capable of reasoning-driven visual tasks, they still exhibit a persistent gap between understanding and generation. The data and code for GIR-Bench are available at https://github.com/HKUST-LongGroup/GIR-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。