arXiv:2604.11632cs.CL2026-04被引 1

评测视觉语言模型对中文艺术的理解与鉴赏能力,发现现有模型在专家级推理上仍有巨大差距。

CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity

论文配图:CArtBench: Evaluating Vision-Language Models on Chinese Art Understanding, Interpretation, and Authenticity
图 1 · 摘自论文原文
  • 构建四类任务评估模型对书画作品的证据推理与风格判断能力
  • 模型在复杂证据链和朝代风格推断上准确率骤降,真伪判别接近随机水平
  • 适合关注艺术智能、跨模态理解及文化认知计算的研究者

我们提出CArtBench,一个基于故宫文物数据的中文艺术理解评测基准,超越简单的识别与问答。该基准包含四项子任务:CURATORQA(基于证据的识别与推理)、CATALOGCAPTION(结构化四段式专家级赏析)、REINTERPRET(可辩护的重新解读并经专家评分)以及CONNOISSEURPAIRS(在视觉相似干扰下的真伪诊断)。数据源自维基数据中带有图像的故宫藏品,并对应权威目录页面,覆盖五个艺术门类及多个朝代。在九个代表性视觉语言模型上测试发现,尽管总体CURATORQA准确率较高,但在关键证据关联和风格-年代推断等难题上表现显著下滑;长文本赏析仍远未达到专家水平;真伪判别能力接近随机,凸显当前模型在鉴定级推理上的严重不足。

原文摘要 · Abstract (English)

We introduce CARTBENCH, a museum-grounded benchmark for evaluating vision-language models (VLMs) on Chinese artworks beyond short-form recognition and QA. CARTBENCH comprises four subtasks: CURATORQA for evidence-grounded recognition and reasoning, CATALOGCAPTION for structured four-section expert-style appreciation, REINTERPRET for defensible reinterpretation with expert ratings, and CONNOISSEURPAIRS for diagnostic authenticity discrimination under visually similar confounds. CARTBENCH is built by aligning image-bearing Palace Museum objects from Wikidata with authoritative catalog pages, spanning five art categories across multiple dynasties. Across nine representative VLMs, we find that high overall CURATORQA accuracy can mask sharp drops on hard evidence linking and style-to-period inference; long-form appreciation remains far from expert references; and authenticity-oriented diagnostic discrimination stays near chance, underscoring the difficulty of connoisseur-level reasoning for current models.

艺术理解视觉语言模型中文多模态真伪鉴别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。