arXiv:2606.22608cs.CVcs.CL2026-06被引 1

构建最大规模楔形文字符号数据集,实现自动识别与文本结构重建。

Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline

论文配图:Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline
图 1 · 摘自论文原文
  • 基于DETR框架,分173/106类进行符号检测,结合排版与词频分析
  • 在eBL数据库中处理8.7万片残片,生成近290万次符号检测结果
  • 无需语言先验,适合大规模古籍数字化与多模态模型接入

学习解读楔形文字泥板是一项极高难度的任务;因此,在约50万片已出土泥板中,仅有极小部分被亚述学家分析。计算机视觉为破译提供了可能,但需大规模密集标注数据集。为此,本文构建迄今最大的标注楔形文字符号数据集,并评估基于可变形检测变压器(DETR)的物体检测模型,采用173类和106类两种分类粒度。所提系统整合自动泥板侧提取、启发式行分组及n-gram文本相似性评估,以连接视觉符号检测与文本结构,相比以往工作在COCO风格检测指标上提升28-37%。推理阶段应用于电子巴比伦图书馆(eBL)语料库中的87,668片泥板残片,生成近290万次符号检测。尽管该方法不依赖语言先验且对泥板损毁与版式变化敏感,仍为全语料库楔形文字分析提供可扩展、可解释的基础,并支持未来与多模态及语言建模框架融合。

原文摘要 · Abstract (English)

Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriologists. Computer vision offers a promising avenue for decipherment but requires large, densely annotated datasets. To address this limitation, the largest annotated cuneiform sign dataset to date is used, and a Deformable Detection Transformer (DETR)-based object detection model is evaluated under two class granularities of 173 and 106 classes. The proposed system integrates automatic tablet-side extraction, heuristic line grouping, and n-gram-based textual similarity evaluation to bridge visual sign detection and textual structure, and achieves consistent improvements of up to 28-37% over prior work on COCO-style detection metrics. At inference, the method is applied to 87,668 tablet fragments from the Electronic Babylonian Library (eBL) corpus, producing nearly 2.9 million sign detections. Although the approach operates without linguistic priors and remains sensitive to tablet damage and layout variability, it provides a scalable and interpretable foundation for corpus-wide cuneiform analysis and supports future integration with multimodal and linguistic modelling frameworks.

古文字识别目标检测数字人文DETR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。