构建图文流动插入评测集,检验多模态模型综合理解能力
FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
- 设计图文流插入任务,要求模型同步理解文本、图像和长篇指令
- 基于625篇新闻构建中英文双语数据集,覆盖10个领域
- 发现顶尖模型如GPT-4o在该任务上仍表现困难,凸显评估挑战
得益于大语言模型(LLMs)和基础视觉模型的突破,大视觉语言模型(LVLMs)取得了显著进展。然而,现有评测基准仅关注单一能力(如识别、检测、理解),难以全面反映其在复杂场景中的潜力。为此,我们提出更具挑战性的“图文流插入任务”(FTII),要求模型同时具备出色的图像理解、指令解析和长文本处理能力。具体而言,在逐步生成的文本段落中,需从候选图像中选出最匹配的一张插入对应段落后。为解决图文序列构建难题,我们采用专业新闻报道作为天然标注标准,构建了包含318篇高质量中文与307篇高质量英文新闻文章的评测集(FTII-Bench),覆盖10个新闻领域。基于此,我们设计两种题型及多级难度。进一步建立基于CLIP和现有LVLMs的双重评估流程,对9个开源、2个闭源LVLM及2个CLIP模型进行评测。结果表明,即使最先进的模型(如GPT-4o)在该任务中也面临显著挑战。
原文摘要 · Abstract (English)
Benefiting from the revolutionary advances in large language models (LLMs) and foundational vision models, large vision-language models (LVLMs) have also made significant progress. However, current benchmarks focus on tasks that evaluating only a single aspect of LVLM capabilities (e.g., recognition, detection, understanding). These tasks fail to fully demonstrate LVLMs' potential in complex application scenarios. To comprehensively assess the performance of existing LVLMs, we propose a more challenging task called the Flow Text with Image Insertion task (FTII). This task requires LVLMs to simultaneously possess outstanding abilities in image comprehension, instruction understanding, and long-text interpretation. Specifically, given several text paragraphs and a set of candidate images, as the text paragraphs accumulate, the LVLMs are required to select the most suitable image from the candidates to insert after the corresponding paragraph. Constructing a benchmark for such a task is highly challenging, particularly in determining the sequence of flowing text and images. To address this challenge, we turn to professional news reports, which naturally contain a gold standard for image-text sequences. Based on this, we introduce the Flow Text with Image Insertion Benchmark (FTII-Bench), which includes 318 high-quality Chinese image-text news articles and 307 high-quality English image-text news articles, covering 10 different news domains. Using these 625 high-quality articles, we construct problems of two different types with multiple levels of difficulty. Furthermore, we establish two different evaluation pipelines based on the CLIP model and existing LVLMs. We evaluate 9 open-source and 2 closed-source LVLMs as well as 2 CLIP-based models. Results indicate that even the most advanced models (e.g., GPT-4o) face significant challenges when tackling the FTII task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。