arXiv:2505.11275cs.MMcs.AI2025-05被引 6

评测大模型对传统文化的理解能力,发现现有模型表现不佳。

TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs

  • 构建双语图文问答基准,聚焦传统中国文化理解。
  • 多模态模型在文化相关视觉推理任务中准确率普遍低于40%。
  • 适合研究跨文化AI、多模态理解的学者与开发者参考。

多模态大语言模型(MLLMs)在多模态内容理解与生成方面取得显著进展,但在非西方文化背景下的应用效果有限,引发对其普适性的担忧。为此,我们提出传统中国文化理解基准(TCC-Bench),一个双语(中文与英文)的图文问答(VQA)基准,专门用于评估MLLMs对传统中国文化的理解能力。TCC-Bench包含来自博物馆文物、日常生活场景、漫画等文化语境的丰富视觉数据。采用半自动化流程:先用GPT-4o纯文本模式生成候选问题,再经人工筛选以保证质量并防止数据泄露。基准设计避免语言偏见,不直接在问题中暴露文化概念。在多种MLLM上的实验表明,当前模型在基于文化背景的视觉推理任务中仍面临严峻挑战。结果凸显了发展更具文化包容性与情境感知能力的多模态系统的重要性。代码与数据详见:https://tcc-bench.github.io/。

原文摘要 · Abstract (English)

Recent progress in Multimodal Large Language Models (MLLMs) have significantly enhanced the ability of artificial intelligence systems to understand and generate multimodal content. However, these models often exhibit limited effectiveness when applied to non-Western cultural contexts, which raises concerns about their wider applicability. To address this limitation, we propose the Traditional Chinese Culture understanding Benchmark (TCC-Bench), a bilingual (i.e., Chinese and English) Visual Question Answering (VQA) benchmark specifically designed for assessing the understanding of traditional Chinese culture by MLLMs. TCC-Bench comprises culturally rich and visually diverse data, incorporating images from museum artifacts, everyday life scenes, comics, and other culturally significant contexts. We adopt a semi-automated pipeline that utilizes GPT-4o in text-only mode to generate candidate questions, followed by human curation to ensure data quality and avoid potential data leakage. The benchmark also avoids language bias by preventing direct disclosure of cultural concepts within question texts. Experimental evaluations across a wide range of MLLMs demonstrate that current models still face significant challenges when reasoning about culturally grounded visual content. The results highlight the need for further research in developing culturally inclusive and context-aware multimodal systems. The code and data can be found at: https://tcc-bench.github.io/.

多模态文化理解评测基准中文AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。