arXiv:2502.15812cs.LGcs.AI2025-02被引 5

首个中文多层级隐含视觉语义评测基准,评估模型对讽刺等深层含义的理解能力。

InsightVision: A Comprehensive, Multi-Level Chinese-based Benchmark for Evaluating Implicit Visual Semantics in Large Vision Language Models

  • 构建四层任务体系:从表层内容到隐含意义逐级评估
  • 15个开源模型与GPT-4o测试显示,最优模型仍落后人类14%
  • 适用于评估中文场景下视觉语言模型的深层语义理解能力

在多模态语言模型发展背景下,理解图像中通过视觉线索传达的微妙含义(如讽刺、侮辱或批评)仍是重大挑战。现有评测基准多集中于直接任务(如图像描述),或仅覆盖幽默、讽刺等有限类别。为此,我们首次提出一个全面、多层级的中文基准——InsightVision,专用于评估大视觉语言模型对隐含意义的理解。该基准系统分为四个子任务:表面内容理解、象征意义解析、背景知识掌握和隐含意义领会。我们提出一种创新的半自动数据构建方法,遵循既定协议。基于此基准,我们评估了15个开源大视觉语言模型及GPT-4o,发现即使表现最佳的模型,在隐含意义理解上仍比人类低近14%。结果凸显当前大视觉语言模型在把握复杂视觉语义方面的根本性局限,为未来研究提供重要方向。论文接受后将公开发布InsightVision数据集与代码。

原文摘要 · Abstract (English)

In the evolving landscape of multimodal language models, understanding the nuanced meanings conveyed through visual cues - such as satire, insult, or critique - remains a significant challenge. Existing evaluation benchmarks primarily focus on direct tasks like image captioning or are limited to a narrow set of categories, such as humor or satire, for deep semantic understanding. To address this gap, we introduce, for the first time, a comprehensive, multi-level Chinese-based benchmark designed specifically for evaluating the understanding of implicit meanings in images. This benchmark is systematically categorized into four subtasks: surface-level content understanding, symbolic meaning interpretation, background knowledge comprehension, and implicit meaning comprehension. We propose an innovative semi-automatic method for constructing datasets, adhering to established construction protocols. Using this benchmark, we evaluate 15 open-source large vision language models (LVLMs) and GPT-4o, revealing that even the best-performing model lags behind human performance by nearly 14% in understanding implicit meaning. Our findings underscore the intrinsic challenges current LVLMs face in grasping nuanced visual semantics, highlighting significant opportunities for future research and development in this domain. We will publicly release our InsightVision dataset, code upon acceptance of the paper.

视觉语言模型隐含语义中文评测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。