构建中文图像隐含意义理解基准,评估模型对中华文化的深层理解能力
Can MLLMs Understand the Deep Implication Behind Chinese Images?
- 基于真实中文网络图片与人工标注,构建文化语境真实可用的评测集
- 模型最高准确率64.4%,远低于人类平均78.2%及峰值81.0%
- 加入情绪提示可提升性能,凸显文化知识短板与改进方向
随着多模态大模型能力持续提升,对其高阶感知与理解能力的评估需求日益迫切。然而,现有工作缺乏针对中文视觉内容高阶理解的评测体系。为此,我们提出CII-Bench——一个旨在评估MLLM对中文图像深层语义理解能力的基准。CII-Bench通过从中国互联网获取真实图像并经人工审核,确保语境真实性;同时包含大量中国传统绘画等文化图像,以检验模型对中国传统文化的理解深度。在多个MLLM上进行的广泛实验显示:模型最高准确率为64.4%,而人类平均达78.2%,峰值高达81.0%,差距显著;模型在传统艺术图像上的表现更差,表明其在高阶语义理解和文化知识储备方面存在明显不足;此外,当在提示中加入图像情绪线索时,多数模型性能有所提升。我们认为,CII-Bench将推动模型更好地理解中文语义与本土化视觉内容,助力迈向专家级通用人工智能。
原文摘要 · Abstract (English)
As the capabilities of Multimodal Large Language Models (MLLMs) continue to improve, the need for higher-order capability evaluation of MLLMs is increasing. However, there is a lack of work evaluating MLLM for higher-order perception and understanding of Chinese visual content. To fill the gap, we introduce the **C**hinese **I**mage **I**mplication understanding **Bench**mark, **CII-Bench**, which aims to assess the higher-order perception and understanding capabilities of MLLMs for Chinese images. CII-Bench stands out in several ways compared to existing benchmarks. Firstly, to ensure the authenticity of the Chinese context, images in CII-Bench are sourced from the Chinese Internet and manually reviewed, with corresponding answers also manually crafted. Additionally, CII-Bench incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, which can deeply reflect the model's understanding of Chinese traditional culture. Through extensive experiments on CII-Bench across multiple MLLMs, we have made significant findings. Initially, a substantial gap is observed between the performance of MLLMs and humans on CII-Bench. The highest accuracy of MLLMs attains 64.4%, where as human accuracy averages 78.2%, peaking at an impressive 81.0%. Subsequently, MLLMs perform worse on Chinese traditional culture images, suggesting limitations in their ability to understand high-level semantics and lack a deep knowledge base of Chinese traditional culture. Finally, it is observed that most models exhibit enhanced accuracy when image emotion hints are incorporated into the prompts. We believe that CII-Bench will enable MLLMs to gain a better understanding of Chinese semantics and Chinese-specific images, advancing the journey towards expert artificial general intelligence (AGI). Our project is publicly available at https://cii-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。