arXiv:2609.02573cs.CV2026-09

构建文本图像深度交织的评测基准,检验多模态模型理解复杂交互的能力

Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment

论文配图:Deeply Interleaved Text-Image Contexts for Multimodal LLMs Assessment
图 1 · 摘自论文原文
  • 提出TIC-Bench评测集,聚焦文本与图像深度交织的上下文理解
  • 涵盖2280个问题,覆盖逻辑、时间、空间三类关联,模型表现远低于人类
  • 适合研究多模态推理、跨模态对齐的学者使用

当前多模态模型的评估与训练主要集中在多图任务上,忽视了文本与图像交错的场景。在这些任务中,文本通常仅作为指令,缺乏与视觉内容的深层语义交互。而真实应用如图文共创、角色追踪和空间重建,要求文本与图像持续互动。因此,模型需具备对深度交织上下文的理解能力。为此,我们提出新评测基准TIC-Bench(深度交织文本-图像上下文),用于评估模型整合文本-图像线索并还原真实事实的能力。该基准包含逻辑、时间、空间三类核心领域,细分为八种类型,共2280个问题。我们评估了10个最先进的多模态大模型,发现其性能显著低于人类专家,且在整合分散于交错视觉与文本输入中的证据方面仍存在持久困难。TIC-Bench为分析和提升多模态模型在深度交织上下文中融合图文信息的能力提供了有效工具。数据集已公开于https://huggingface.co/datasets/pino10010/TIC-Bench。

原文摘要 · Abstract (English)

Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applications like text-image co-creation, character tracking, and spatial reconstruction require constant interaction between text and images. Consequently, models must possess a deep understanding of these interleaved contexts. To bridge this gap, we introduce a novel benchmark, TIC-Bench (deeply interleaved Text-Image Contexts), designed to evaluate the capability of models to integrate text-image clues and recover the ground truth facts within deeply interleaved contexts. This benchmark encompasses three core domains: Logical, Temporal, and Spatial Association, which are further categorized into eight specific types, comprising a total of 2,280 questions. We evaluated 10 state-of-the-art MLLMs and observed a substantial performance gap compared to human experts, together with persistent difficulties in integrating evidence distributed across interleaved visual and textual inputs. Ultimately, this benchmark provides a valuable analytical tool for assessing and advancing the ability of multimodal models to effectively integrate text and image information in deeply interleaved contexts. TIC-Bench is publicly available at https://huggingface.co/datasets/pino10010/TIC-Bench

多模态评测基准图文理解上下文融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。