构建2万张甲骨文图像数据集,评估大模型视觉解码能力。
PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters
- 构建含2万张甲骨文与实物图像的多选题数据集
- 大模型依赖语言先验,未有效利用视觉信息
- 适合研究视觉-语言模型在古文字识别中的应用
破解甲骨文(OBC),作为最早可考证的汉字形式,是学者们长期追求的目标,对理解人类早期生产方式具有不可替代的意义。当前甲骨文破译方法受限于考古发掘的零散性及铭文语料有限。借助大视觉-语言模型(LMMs)强大的视觉感知能力,其在甲骨文视觉解码方面的潜力日益显现。本文提出PictOBI-20k,一个专用于评估LMMs在象形甲骨文视觉解码任务中表现的数据集,包含20,000张精心收集的甲骨文与真实物体图像,形成超过15,000个多选题。我们还进行主观标注,探究人类与LMMs在视觉推理中的参考点一致性。实验表明,通用大模型具备初步的视觉解码能力,但大多情况下受语言先验制约,未能有效利用视觉信息。我们希望该数据集能推动未来面向甲骨文的LMMs在视觉注意力方面的评估与优化。代码与数据集将公开于https://github.com/OBI-Future/PictOBI-20k。
原文摘要 · Abstract (English)
Deciphering oracle bone characters (OBCs), the oldest attested form of written Chinese, has remained the ultimate, unwavering goal of scholars, offering an irreplaceable key to understanding humanity's early modes of production. Current decipherment methodologies of OBC are primarily constrained by the sporadic nature of archaeological excavations and the limited corpus of inscriptions. With the powerful visual perception capability of large multimodal models (LMMs), the potential of using LMMs for visually deciphering OBCs has increased. In this paper, we introduce PictOBI-20k, a dataset designed to evaluate LMMs on the visual decipherment tasks of pictographic OBCs. It includes 20k meticulously collected OBC and real object images, forming over 15k multi-choice questions. We also conduct subjective annotations to investigate the consistency of the reference point between humans and LMMs in visual reasoning. Experiments indicate that general LMMs possess preliminary visual decipherment skills, and LMMs are not effectively using visual information, while most of the time they are limited by language priors. We hope that our dataset can facilitate the evaluation and optimization of visual attention in future OBC-oriented LMMs. The code and dataset will be available at https://github.com/OBI-Future/PictOBI-20k.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。