构建中文文化遗产多模态理解评测基准,聚焦推理过程质量。
CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage

- 设计包含50项任务的跨博物馆多模态评测集
- 提出自适应维度加权的推理评分方法ReaScore
- 揭示主流模型答案准确但推理过程不完整
评估多模态大模型在中文文化遗产(CCH)中的理解能力,需对视觉、文本、风格和历史线索进行细粒度推理。现有CCH评测多关注最终答案准确性,而推理过程的准确性和完整性未受重视。为此,我们提出CulMind与CulMind-R:一个涵盖100+博物馆、覆盖50个任务的高质量多模态CCH评测基准,以及24个任务的推理子集,可自适应定义任务相关的推理维度。为评估推理质量,我们提出ReaScore——一种任务自适应的评分指标,通过自动加权任务相关维度评估推理。在14个领先多模态大模型上的实验显示,答案与推理间存在显著差距,尤其在挑战性任务上。进一步分析表明,任务自适应维度选择与加权更符合专家判断。整体上,该基准与评估方法支持更贴近专家认知的遗产理解评估,并为文化传承领域评估提供可迁移范式。数据、代码与评估脚本已公开于https://github.com/ZevTsao/CulMind。
原文摘要 · Abstract (English)
Evaluating Multimodal Large Language Models (MLLMs) in Chinese Cultural Heritage (CCH) requires fine-grained reasoning over visual, textual, stylistic, and historical clues. However, existing CCH benchmarks mainly emphasize final-answer accuracy, while the accuracy and completeness of reasoning processes remain underexplored. To address this gap, we introduce CulMind and CulMind-R: a high-quality benchmark for multimodal CCH covering 50 tasks from collections of more than 100 museums, and a 24-task reasoning subset that adaptively defines task-specific dimensions for reasoning process evaluation. To evaluate reasoning quality, we propose ReaScore, a task-adaptive metric that evaluates reasoning by automatically weighting task-relevant dimensions. Experiments on 14 leading MLLMs reveal a substantial gap between answers and reasoning, especially on challenging tasks. Further analysis shows that task-adaptive dimension selection and weighting better align evaluation results with expert judgments. Overall, our benchmark and metric support a more expert-aligned assessment of CCH understanding and offer a transferable reference for broader evaluations of cultural heritage. We publicly release the data, code, and evaluation scripts at https://github.com/ZevTsao/CulMind to facilitate reproducible research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。