构建首个句子级甲骨文理解评测基准,检验大模型是否超越单字识别。
Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

- 用标准字形替换原拓片字符,生成清晰句子级甲骨文图像
- 695组问答测试显示当前大模型在句级理解上表现不佳
- 揭示视觉错误会沿推理链传播,说明仍依赖字级识别
现有甲骨文视觉识别研究多聚焦单字级别,忽视完整卜辞中的长文本连贯性与上下文依赖。近年来多模态大模型(MLLM)强大的视觉感知能力为甲骨文信息处理带来新可能。本文提出S-OBI,一个面向句子级甲骨文理解的新型评测基准。不同于使用模糊不全的拓片作为输入,S-OBI通过字形替换与组合,合成清晰标准化的句子级甲骨文实例。基于95份经专家鉴定、修正和验证的原始拓片及其译文,将原拓片中的字符替换为来自现有甲骨文数据集的干净字形样本,同时保留整体刻写结构与语义组织。该方法削弱低层失真影响,实现更聚焦的句级理解评估。在此基础上,设计语义匹配、语义槽抽取与上下文推理任务,构建695组问答对。实验表明,当前主流MLLM在句子级甲骨文理解上表现较差,尤其未遮挡区域的视觉识别错误会通过推理链传播,导致遮挡字识别出错,说明当前模型在句级理解上仍严重依赖字级识别能力。S-OBI为诊断模型能否从孤立字识别迈向结构化卜辞理解提供有效工具。
原文摘要 · Abstract (English)
Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contextual dependencies embedded in complete divination charges. Recently, the powerful visual perception capabilities of multimodal large language models (MLLMs) have opened new possibilities for OBI information processing. In this work, we introduce S-OBI, a novel benchmark for evaluating MLLMs in Sentence-level OBI understanding. Instead of using noisy and incomplete rubbings as the visual input, S-OBI synthesizes clear and standardized sentence-level OBI instances through glyph substitution and composition. According to 95 original rubbings with translations that have been identified, corrected, and verified by experts, we replace characters in the original rubbings with corresponding clean glyph samples sourced from existing OBI datasets while preserving the overall inscriptional structure and semantic organization. This mitigates the influence of low-level distortions and enables a more focused evaluation of sentence-level OBI understanding. Based on this, we design semantic matching, semantic slot extraction, and contextual reasoning tasks and obtain 695 question-answer pairs. Experiments reveal the inferiority of contemporary MLLMs on sentence-level OBI understanding. In particular, visual perception errors in unmasked regions propagate through the reasoning chain, leading to erroneous predictions for masked characters, which indicates that sentence-level OBI understanding in current models remains strongly dependent on character-level recognition. Overall, S-OBI provides a diagnostic benchmark for evaluating whether MLLMs can move beyond isolated character recognition toward structured inscription-level understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。