用漫画构建跨文化多模态评测基准,挑战大模型文化理解能力
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
- 基于漫画设计跨文化多任务评测,覆盖从视觉识别到文化冲突理解
- 包含2000+图像和18000+问答对,涵盖三类渐进难度任务
- 测试11个开源模型,揭示大模型与人类在文化理解上的巨大差距
文化意识已成为多模态大语言模型(MLLMs)的关键能力。然而现有评测基准在任务设计上难度不足,缺乏跨语言任务,且多使用真实图像——每张图通常只体现单一文化,导致评测过于简单。为此,我们提出C³B(Comics Cross-Cultural Benchmark),一个新型的多文化、多任务、多语言文化意识评测基准。C³B包含2000多张图像和18000多个问答对,涵盖三类渐进难度任务:从基础视觉识别,到高层次文化冲突理解,最终到文化内容生成。我们在11个开源MLLM上进行了评估,结果显示模型表现与人类水平存在显著差距,表明C³B对当前MLLM构成实质性挑战,推动未来研究提升模型的文化认知能力。
原文摘要 · Abstract (English)
Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their task design and are deficient in cross-lingual tasks. Moreover, current benchmarks often use real-world images. Each real-world image typically contains one culture, making these benchmarks relatively easy for MLLMs. Based on this, we propose C$^3$B (Comics Cross-Cultural Benchmark), a novel multicultural, multitask and multilingual cultural awareness capabilities benchmark. C$^3$B comprises over 2000 images and over 18000 QA pairs, constructed on three tasks with progressed difficulties, from basic visual recognition to higher-level cultural conflict understanding, and finally to cultural content generation. We conducted evaluations on 11 open-source MLLMs, revealing a significant performance gap between MLLMs and human performance. The gap demonstrates that C$^3$B poses substantial challenges for current MLLMs, encouraging future research to advance the cultural awareness capabilities of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。