用跨模态融合解析甲骨文,提升古文字识别准确率
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
- 分层视觉-语义模块+图结构推理,结合多粒度特征与语义关系
- 在新数据集OracleSem上,识别准确率超越现有模型23.6个百分点
- 适合古文字研究、文化遗产数字化及跨模态理解方向学者
甲骨文作为中国最早的成熟文字系统,因复杂的象形结构和与现代汉字的差异,在自动识别方面面临巨大挑战。本文提出OracleSage,一种融合分层视觉理解与基于图的语义推理的跨模态框架。具体包括:(1) 分层视觉-语义理解模块,通过逐步微调LLaVA的视觉主干网络实现多粒度特征提取;(2) 基于图的语义推理框架,利用动态消息传递捕捉视觉组件与语义概念之间的关系;(3) OracleSem,一个包含完整象形与语义标注的丰富甲骨文数据集。实验表明,OracleSage显著优于当前最先进的视觉-语言模型。该研究为古代文本解读建立了新范式,并为考古学研究提供了关键技术支撑。
原文摘要 · Abstract (English)
Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We introduce OracleSage, a novel cross-modal framework that integrates hierarchical visual understanding with graph-based semantic reasoning. Specifically, we propose (1) a Hierarchical Visual-Semantic Understanding module that enables multi-granularity feature extraction through progressive fine-tuning of LLaVA's visual backbone, (2) a Graph-based Semantic Reasoning Framework that captures relationships between visual components and semantic concepts through dynamic message passing, and (3) OracleSem, a semantically enriched OBS dataset with comprehensive pictographic and semantic annotations. Experimental results demonstrate that OracleSage significantly outperforms state-of-the-art vision-language models. This research establishes a new paradigm for ancient text interpretation while providing valuable technical support for archaeological studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。