arXiv:2604.06711cs.CVcs.CL2026-04ACL被引 1

用部件知识增强大模型,让甲骨文识别更准确。

Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation

  • 结合视觉与语言模型,自动分析甲骨文字形部件
  • 在三个任务上表现优于基线,准确率显著提升
  • 适合古文字研究者与跨模态智能方向开发者

解读古代中国甲骨文是一项具有挑战性的任务,可揭示古代信仰、制度与文化。现有方法将其视为封闭集图像识别,难以弥合“解读鸿沟”:单个字符虽罕见独特,却由有限的、具象化的重复部件构成,这些部件携带可迁移的语义。为此,我们提出一种代理驱动的视觉-语言模型框架,融合精准视觉定位的视觉-语言模型与基于大语言模型的代理,实现部件识别、基于图的知识检索与关系推理的自动化推理链,以获得语言学上准确的解读。为支持该框架,我们构建了OB-Radix数据集,该数据集由专家标注,包含1,022张字符图像(934个唯一字符)和1,853张细粒度部件图像,涵盖478个不同部件及其经验证的解释,弥补了以往语料库的结构性与语义性缺失。在三个不同任务的基准测试中,我们的系统相较基线方法生成了更详尽、更精确的解读结果。

原文摘要 · Abstract (English)

Deciphering ancient Chinese Oracle Bone Script (OBS) is a challenging task that offers insights into the beliefs, systems, and culture of the ancient era. Existing approaches treat decipherment as a closed-set image recognition problem, which fails to bridge the ``interpretation gap'': while individual characters are often unique and rare, they are composed of a limited set of recurring, pictographic components that carry transferable semantic meanings. To leverage this structural logic, we propose an agent-driven Vision-Language Model (VLM) framework that integrates a VLM for precise visual grounding with an LLM-based agent to automate a reasoning chain of component identification, graph-based knowledge retrieval, and relationship inference for linguistically accurate interpretation. To support this, we also introduce OB-Radix, an expert-annotated dataset providing structural and semantic data absent from prior corpora, comprising 1,022 character images (934 unique characters) and 1,853 fine-grained component images across 478 distinct components with verified explanations. By evaluating our system across three benchmarks of different tasks, we demonstrate that our framework yields more detailed and precise decipherments compared to baseline methods.

甲骨文多模态知识增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。