评测多图上下文生成能力,提出新方法提升图像一致性。
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
- 构建六任务基准测试多图生成与推理能力。
- 新方法DAR在不训练情况下改善生成连贯性。
- 适合研究多模态生成与模型评估的学者使用。
统一多模态模型(UMMs)在图像理解与生成方面取得显著进展,但现有评测大多聚焦文本到图像或单图编辑,忽视多图上下文生成挑战。本文提出MICON-Bench,涵盖六项任务,评估跨图组合、上下文推理与身份一致性。引入基于多模态大语言模型(MLLM)的按检查点自动验证框架,实现语义与视觉一致性检验。进一步提出无需训练的动态注意力重平衡(DAR)机制,在推理时动态调整注意力以增强连贯性、减少幻觉。在多个开源先进模型上实验表明,MICON-Bench能有效暴露多图推理缺陷,DAR显著提升生成质量与跨图一致性。代码已开源。
原文摘要 · Abstract (English)
Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related images, existing benchmarks rarely address the challenges of multi-image context generation, focusing mainly on text-to-image or single-image editing tasks. In this work, we introduce \textbf{MICON-Bench}, a comprehensive benchmark covering six tasks that evaluate cross-image composition, contextual reasoning, and identity preservation. We further propose an MLLM-driven Evaluation-by-Checkpoint framework for automatic verification of semantic and visual consistency, where multimodal large language model (MLLM) serves as a verifier. Additionally, we present \textbf{Dynamic Attention Rebalancing (DAR)}, a training-free, plug-and-play mechanism that dynamically adjusts attention during inference to enhance coherence and reduce hallucinations. Extensive experiments on various state-of-the-art open-source models demonstrate both the rigor of MICON-Bench in exposing multi-image reasoning challenges and the efficacy of DAR in improving generation quality and cross-image coherence. Github: https://github.com/Angusliuuu/MICON-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。