构建多轮图文交互评测基准,揭示大模型生成缺陷与改进方法
IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

- 设计三类动态多轮图文对话数据集,覆盖3113个样本、12034次交互
- 发现主流模型在多轮对话中存在严重生成偏差,准确率显著下降
- 验证思维链等策略可有效提升生成质量,适合研究多模态交互的团队
近年来,统一多模态模型(UMMs)在单一框架内同时支持理解与生成。在真实应用中,掌握动态、多轮交替的图文对话是关键任务。然而,现有基准多局限于单轮或静态设置,且忽视多轮交互中的暴露偏差。为此,我们提出IMUG-Bench,一个全面评估多轮交替图文对话中理解与生成能力的基准。该基准包含静态空间、时序因果和混合三类,共3,113个样本、12,034次交互回合,并引入动态理解问题,更贴近真实多轮交互场景。大规模实验系统评估主流开源与闭源UMMs,揭示其能力边界与失效模式,发现生成侧存在显著暴露偏差。我们进一步探索链式思考、自验证和Best-of-N采样等测试时扩展策略,显著提升生成准确率并缓解生成偏差。这些发现为提升未来UMMs的鲁棒性与多轮交互能力提供重要启示。
原文摘要 · Abstract (English)
In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework. Mastering dynamic, multi-turn interleaved image-text dialogues is a crucial task for UMMs in real-world applications. However, existing benchmarks fail to evaluate this important task, as they are often limited to single-turn or static settings, and typically overlook exposure bias in multi-turn interactions. To bridge this gap, we propose IMUG-Bench, a comprehensive benchmark for multi-turn interleaved image-text dialogue of UMMs that jointly evaluates their understanding and generation capabilities. Our IMUG-Bench comprises three classes: Static Spatial, Temporal Causal, and Hybrid, covering 3,113 samples and 12,034 interaction turns. It also includes dynamic understanding questions, thereby supporting evaluation that better reflects real-world multi-turn interaction scenarios. Large-scale experiments on IMUG-Bench systematically evaluate mainstream open-source and closed-source UMMs, revealing their capability boundaries and failure modes, and uncovering pronounced exposure bias on the generation side in multi-turn interactions. We further explore several test-time scaling strategies, including Chain-of-Thought, Self-Verification, and Best-of-N Sampling, which effectively improve generation accuracy and mitigate exposure bias in generation tasks. These findings provide insights into enhancing the robustness and multi-turn interaction capability of future UMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。