用生成字典检索破解甲骨文,准确率超50%。
Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval
- 将破译任务转为字典检索,基于字符演化生成合成字典。
- 对未见字符的Top-10准确率达54.3%,Top-50达86.6%。
- 结果可解释,适合考古与古文字研究者使用。
理解人类最早的书写系统对重建文明起源至关重要,但许多古代文字仍未被破译。中国商代的甲骨文(OBS)即为典型:约4600个字符中仅1500个被解读,大量3000年前的铭文仍部分未知。受限于极端数据稀缺,现有计算方法在未见字符上的准确率低于3%。本文将破译从分类任务重构为基于字典的检索任务,利用遵循字符演化规律的深度学习生成现代汉字对应的可能性甲骨文变体合成字典。学者可通过查询未知铭文,检索视觉相似的候选字并获得透明证据,取代算法黑箱。该方法对未见字符的Top-10准确率达54.3%,Top-50达86.6%。此可扩展、可解释的框架加速了关键未解文字系统的破译,并为人工智能辅助考古发现提供通用方法。
原文摘要 · Abstract (English)
Understanding humanity's earliest writing systems is crucial for reconstructing civilization's origins, yet many ancient scripts remain undeciphered. Oracle Bone Script (OBS) from China's Shang dynasty exemplifies this challenge: only approximately 1,500 of roughly 4,600 characters have been decoded, and a substantial portion of these 3,000-year-old inscriptions remains only partially understood. Limited by extreme data scarcity, existing computational methods achieve under 3% accuracy on unseen characters -- the core palaeographic challenge. We overcome this by reframing decipherment from classification to dictionary-based retrieval. Using deep learning guided by character evolution principles, we generate a comprehensive synthetic dictionary of plausible OBS variants for modern Chinese characters. Scholars query unknown inscriptions to retrieve visually similar candidates with transparent evidence, replacing algorithmic black boxes with interpretable hypotheses. Our approach achieves 54.3% Top-10 and 86.6% Top-50 accuracy for unseen characters. This scalable, transparent framework accelerates decipherment of a pivotal undeciphered script and establishes a generalizable methodology for AI-assisted archaeological discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。