评测大模型识别基督教圣像画的能力,发现部分模型已超越传统方法。
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
- 用类标签、描述词和少量样例三种方式测试多模态大模型
- Gemini-2.5 Pro和GPT-4o在多数数据集上超过ResNet50基线
- 加描述词能提升零样本性能,但少样本学习效果有限
本研究评估了多模态大语言模型(MLLMs)和视觉语言模型(VLMs)在基督教圣像画单标签分类任务中的表现。目标是检验通用型模型(如CLIP、SigLIP、GPT-4o、Gemini 2.5)能否胜任传统由监督分类器处理的图像理解任务。研究基于三个原生支持Iconclass的数据库:ArtDL、ICONCLASS和Wikidata,筛选出前10个最常见类别。模型在三种条件下测试:(1)仅用类别标签,(2)加入Iconclass描述,(3)使用五张示例进行少样本学习。结果与在相同数据集上微调的ResNet50基线对比。结果显示,Gemini-2.5 Pro和GPT-4o在多数数据集上优于基线;在Wikidata上,尽管准确率下降,但SigLIP表现最佳,表明模型对图像尺寸和元数据一致性敏感。添加类别描述普遍提升了零样本性能,而少样本学习仅带来偶发且微弱的准确率提升。结论表明,通用多模态大模型具备处理复杂文化遗产图像分类的能力,可作为数字人文工作流中的元数据标注工具,未来可探索提示优化及更广泛的模型与策略扩展。
原文摘要 · Abstract (English)
This study evaluates the capabilities of Multimodal Large Language Models (LLMs) and Vision Language Models (VLMs) in the task of single-label classification of Christian Iconography. The goal was to assess whether general-purpose VLMs (CLIP and SigLIP) and LLMs, such as GPT-4o and Gemini 2.5, can interpret the Iconography, typically addressed by supervised classifiers, and evaluate their performance. Two research questions guided the analysis: (RQ1) How do multimodal LLMs perform on image classification of Christian saints? And (RQ2), how does performance vary when enriching input with contextual information or few-shot exemplars? We conducted a benchmarking study using three datasets supporting Iconclass natively: ArtDL, ICONCLASS, and Wikidata, filtered to include the top 10 most frequent classes. Models were tested under three conditions: (1) classification using class labels, (2) classification with Iconclass descriptions, and (3) few-shot learning with five exemplars. Results were compared against ResNet50 baselines fine-tuned on the same datasets. The findings show that Gemini-2.5 Pro and GPT-4o outperformed the ResNet50 baselines. Accuracy dropped significantly on the Wikidata dataset, where Siglip reached the highest accuracy score, suggesting model sensitivity to image size and metadata alignment. Enriching prompts with class descriptions generally improved zero-shot performance, while few-shot learning produced lower results, with only occasional and minimal increments in accuracy. We conclude that general-purpose multimodal LLMs are capable of classification in visually complex cultural heritage domains. These results support the application of LLMs as metadata curation tools in digital humanities workflows, suggesting future research on prompt optimization and the expansion of the study to other classification strategies and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。