新基准评估多模态模型在图文场景中的认知能力
MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark
- 设计图文场景认知评测任务,融合视觉推理与内容创作
- 对比模型在感知与认知能力上的差距,揭示认知短板
- 支持自动评估,适合研究多模态认知能力的学者
文本丰富的视觉场景理解已成为评估多模态大语言模型(MLLMs)的重要方向,因其广泛应用前景。现有基准侧重感知能力评估,忽视了认知能力的考察。为此,我们提出MCTBench——一个面向文本丰富视觉场景的多模态认知评测基准,通过视觉推理和内容生成任务评估MLLMs的认知能力。为减少不同数据集分布带来的评估偏差,MCTBench包含多项感知任务(如场景文本识别),确保对模型感知与认知能力的公平比较。同时,构建自动化评估流程以提升内容生成任务的效率与公正性。在MCTBench上对多种MLLMs的评估表明,尽管其感知能力出色,认知能力仍需提升。我们期望MCTBench能为社区提供高效资源,推动文本丰富视觉场景下认知能力的研究与优化。
原文摘要 · Abstract (English)
The comprehension of text-rich visual scenes has become a focal point for evaluating Multi-modal Large Language Models (MLLMs) due to their widespread applications. Current benchmarks tailored to the scenario emphasize perceptual capabilities, while overlooking the assessment of cognitive abilities. To address this limitation, we introduce a Multimodal benchmark towards Text-rich visual scenes, to evaluate the Cognitive capabilities of MLLMs through visual reasoning and content-creation tasks (MCTBench). To mitigate potential evaluation bias from the varying distributions of datasets, MCTBench incorporates several perception tasks (e.g., scene text recognition) to ensure a consistent comparison of both the cognitive and perceptual capabilities of MLLMs. To improve the efficiency and fairness of content-creation evaluation, we conduct an automatic evaluation pipeline. Evaluations of various MLLMs on MCTBench reveal that, despite their impressive perceptual capabilities, their cognition abilities require enhancement. We hope MCTBench will offer the community an efficient resource to explore and enhance cognitive capabilities towards text-rich visual scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。