用视觉语言模型解析甲骨文,结合部件与图像理解提升可解释性。
Interpretable Oracle Bone Script Decipherment through Radical and Pictographic Analysis with LVLMs
- 通过渐进式训练让模型从部件识别到图像语义推理
- 在零样本设置下达到当前最佳的10-准确率
- 提供可追溯的分析过程,适合历史研究者使用
作为最古老成熟的文字系统,甲骨文因稀有、抽象和象形多样性长期难以破译。现有深度学习方法虽取得进展,但常忽略字形与语义间的复杂关联,导致泛化性和可解释性不足,尤其在零样本和未破译字形场景下表现有限。为此,我们提出一种基于大视觉语言模型的可解释甲骨文破译方法,融合部件分析与象形-语义理解,建立字形到意义的推理桥梁。具体地,设计渐进式训练策略,引导模型从部件识别逐步过渡至象形分析与互证;并引入受分析结果启发的部件-象形双重匹配机制,显著提升零样本破译性能。为支持训练,构建了包含47,157个汉字的象形破译甲骨文数据集,附带甲骨文图像与象形分析文本。在公开基准测试中,本方法实现最先进的Top-10准确率,并具备优异的零样本能力。更重要的是,模型输出逻辑清晰的分析过程,或可为未破译甲骨文提供考古学参考,具有数字人文与历史研究潜力。数据集与代码将开源于https://github.com/PKXX1943/PD-OBS。
原文摘要 · Abstract (English)
As the oldest mature writing system, Oracle Bone Script (OBS) has long posed significant challenges for archaeological decipherment due to its rarity, abstractness, and pictographic diversity. Current deep learning-based methods have made exciting progress on the OBS decipherment task, but existing approaches often ignore the intricate connections between glyphs and the semantics of OBS. This results in limited generalization and interpretability, especially when addressing zero-shot settings and undeciphered OBS. To this end, we propose an interpretable OBS decipherment method based on Large Vision-Language Models, which synergistically combines radical analysis and pictograph-semantic understanding to bridge the gap between glyphs and meanings of OBS. Specifically, we propose a progressive training strategy that guides the model from radical recognition and analysis to pictographic analysis and mutual analysis, thus enabling reasoning from glyph to meaning. We also design a Radical-Pictographic Dual Matching mechanism informed by the analysis results, significantly enhancing the model's zero-shot decipherment performance. To facilitate model training, we propose the Pictographic Decipherment OBS Dataset, which comprises 47,157 Chinese characters annotated with OBS images and pictographic analysis texts. Experimental results on public benchmarks demonstrate that our approach achieves state-of-the-art Top-10 accuracy and superior zero-shot decipherment capabilities. More importantly, our model delivers logical analysis processes, possibly providing archaeologically valuable reference results for undeciphered OBS, and thus has potential applications in digital humanities and historical research. The dataset and code will be released in https://github.com/PKXX1943/PD-OBS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。