arXiv:2509.17040cs.CVcs.AI2025-09ICCV被引 10

构建多图交错推理新基准,提升大模型跨模态理解能力

From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning

  • 设计含交错文本的多图推理任务,强化跨模态关联理解
  • 提出从易到难的渐进式训练策略,显著提升模型性能
  • 适用于研究复杂跨模态推理的学者与工程师

多图像交错推理旨在提升多模态大语言模型(MLLMs)对多张图像及其关联文本的联合理解与推理能力,带来超越单图或非交错多图任务的独特挑战。现有基准普遍忽略交错文本上下文,且忽视图像与对应文本间的特定关系。为此,我们提出新型基准MIR,要求模型在多图像及交错文本背景下进行联合推理,准确匹配图像区域与对应文本,并逻辑连接跨图像信息。为提升模型对多图像交错数据的理解能力,我们在基准中引入每例的推理步骤,并提出分阶段渐进式课程学习策略,遵循‘由易到难’原则,逐步引导模型应对复杂场景。大量实验表明,该方法显著提升了多个MLLMs在MIR及其他基准上的推理表现。我们认为MIR将推动多图像交错推理研究,促进MLLMs在复杂跨模态任务中的发展。

原文摘要 · Abstract (English)

Multi-image Interleaved Reasoning aims to improve Multi-modal Large Language Models (MLLMs) ability to jointly comprehend and reason across multiple images and their associated textual contexts, introducing unique challenges beyond single-image or non-interleaved multi-image tasks. While current multi-image benchmarks overlook interleaved textual contexts and neglect distinct relationships between individual images and their associated texts, enabling models to reason over multi-image interleaved data may significantly enhance their comprehension of complex scenes and better capture cross-modal correlations. To bridge this gap, we introduce a novel benchmark MIR, requiring joint reasoning over multiple images accompanied by interleaved textual contexts to accurately associate image regions with corresponding texts and logically connect information across images. To enhance MLLMs ability to comprehend multi-image interleaved data, we introduce reasoning steps for each instance within the benchmark and propose a stage-wise curriculum learning strategy. This strategy follows an "easy to hard" approach, progressively guiding models from simple to complex scenarios, thereby enhancing their ability to handle challenging tasks. Extensive experiments benchmarking multiple MLLMs demonstrate that our method significantly enhances models reasoning performance on MIR and other established benchmarks. We believe that MIR will encourage further research into multi-image interleaved reasoning, facilitating advancements in MLLMs capability to handle complex inter-modal tasks.

多模态推理大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。