构建古文字评测基准,检验大模型在甲骨文全链条处理中的能力。
OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?
- 设计涵盖识别、拼接、分类等五类任务的甲骨文多模态评测集
- 6种商用与17种开源大模型在精细感知任务上仍远逊于人类专家
- 大模型在释读中表现接近未训练人类,展现创意解读潜力
我们提出OBI-Bench,一个系统性评估大模型在甲骨文全链条处理任务中表现的综合性基准,要求具备领域专业知识与深思熟虑能力。该基准包含5,523张来源多样、精心采集的图像,覆盖识别、拼接、分类、检索和释读五大关键问题,涵盖从考古发掘到合成阶段的多阶段字形,如原始甲骨、墨拓、碎片、单字裁剪及手写字符。不同于现有基准,OBI-Bench强调基于甲骨文特有知识的高级视觉感知与推理,挑战模型完成专家级任务。对6种专有大模型及17种开源大模型的评估表明,即便GPT-4o、Gemini 1.5 Pro和Qwen-VL-Max等最新版本,在部分细粒度感知任务上仍显著落后于普通人类。然而,在释读任务中,其表现已接近未受训人类,显示出提供新解释视角和生成创造性推测的潜力。我们期望OBI-Bench能推动领域专用多模态基础模型发展,深化对古代语言研究的探索,并挖掘大模型尚未释放的潜能。
原文摘要 · Abstract (English)
We introduce OBI-Bench, a holistic benchmark crafted to systematically evaluate large multi-modal models (LMMs) on whole-process oracle bone inscriptions (OBI) processing tasks demanding expert-level domain knowledge and deliberate cognition. OBI-Bench includes 5,523 meticulously collected diverse-sourced images, covering five key domain problems: recognition, rejoining, classification, retrieval, and deciphering. These images span centuries of archaeological findings and years of research by front-line scholars, comprising multi-stage font appearances from excavation to synthesis, such as original oracle bone, inked rubbings, oracle bone fragments, cropped single characters, and handprinted characters. Unlike existing benchmarks, OBI-Bench focuses on advanced visual perception and reasoning with OBI-specific knowledge, challenging LMMs to perform tasks akin to those faced by experts. The evaluation of 6 proprietary LMMs as well as 17 open-source LMMs highlights the substantial challenges and demands posed by OBI-Bench. Even the latest versions of GPT-4o, Gemini 1.5 Pro, and Qwen-VL-Max are still far from public-level humans in some fine-grained perception tasks. However, they perform at a level comparable to untrained humans in deciphering tasks, indicating remarkable capabilities in offering new interpretative perspectives and generating creative guesses. We hope OBI-Bench can facilitate the community to develop domain-specific multi-modal foundation models towards ancient language research and delve deeper to discover and enhance these untapped potentials of LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。