arXiv:2604.27389cs.CVcs.AI2026-04

评测大模型在图文混排场景下的精细对齐能力

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts

论文配图:COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts
图 1 · 摘自论文原文
  • 构建图文交错场景的细粒度对齐评测基准
  • 覆盖4个领域,含6161道高质量问题
  • 提供6类错误分析,定位模型短板

近年来,多模态大语言模型(MLLMs)在多项多模态评测中取得显著进展。然而,现有评测主要集中在单图或多图理解上。在文档阅读等真实场景中,信息常以图文交错形式呈现,要求模型不仅能识别图像内容,还需在交错上下文中精准定位文本与视觉证据,建立细粒度对齐,并基于上下文推理。目前缺乏系统性基准来量化MLLM在交错图文场景中的细粒度理解能力。为此,我们提出COHERENCE,一个用于评估MLLM恢复交错多模态上下文中细粒度图文对应关系的能力的基准。COHERENCE涵盖四个代表性领域的交错图文内容,包含6,161道高质量问题。此外,我们进行六类错误分析,实现对当前MLLM在交错图文理解中失败原因的细粒度归因。

原文摘要 · Abstract (English)

In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these advances, most existing benchmarks mainly focus on single-image or multi-image comprehension. In real-world scenarios such as document reading, information is often presented as interleaved multimodel contexts. This requires MLLMs not only to recognize the content of individual images, but also to identify relevant textual and visual evidence, establish fine-grained alignments between them, and reason over these aligned signals in interleaved contexts based on contextual evidence. However, there is still a lack of systematic benchmarks for quantifying the fine-grained understanding ability of MLLMs in interleaved image-text contexts. To fill this gap, we propose COHERENCE, a benchmark designed to evaluate the ability of MLLMs to recover fine-grained image-text correspondences in interleaved multimodal contexts. COHERENCE covers interleaved image-text content from four representative domains and contains 6,161 high-quality questions. Moreover, we perform a six-type error analysis, enabling fine-grained attribution of failures in interleaved image-text understanding to the specific capabilities missing in current MLLMs.

多模态图文对齐评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。