arXiv:2608.30616cs.CV2026-08

用少量标注数据让OCR读懂手写陶器记录,实现自动结构化提取。

OCR-Based Field Extraction for Archaeological Pottery Metadata: The CENTURIA Dataset

论文配图:OCR-Based Field Extraction for Archaeological Pottery Metadata: The CENTURIA Dataset
图 1 · 摘自论文原文
  • 用LoRA微调57个样本,小数据训练提升手写文本识别准确率。
  • 微调后错误率低于1.5%,字段级准确率达87%以上。
  • 适合需要处理历史手稿的考古、档案数字化研究者。

陶器是重建古代社会时间与经济格局的关键资料。考古学家常通过技术绘图和手写元数据记录陶器发现,这些信息对断代、来源判定和跨遗址比较至关重要,但因难以被计算分析,需逐条手动转录。本文研究先进文档分析模型在此任务中的表现,提出CENTURIA数据集,包含来自罗马时期卡鲁恩图姆遗址的507条陶器记录,涵盖7个元数据类别,提供文本转录、边界框及字段级标签。五种OCR模型的基准测试显示显著领域差距:零样本转录错误率达15-32% SpACER-M,远高于印刷文档,且特定领域字段识别率不足3%。仅用57个样本进行LoRA微调(符合实际注释预算),即可将错误率降至1.5%以下,整体字段准确率超87%。结果表明,少量专家验证的微调数据足以将手写陶器记录转化为可搜索的结构化元数据,适用于考古数据库。

原文摘要 · Abstract (English)

Pottery is a primary source for reconstructing the chronological and economic dimensions of past societies. Archaeologists often document ceramic finds through technical drawings and handwritten metadata. This metadata is critical for dating, provenance attribution, and cross-site comparison, but remains inaccessible to computational analysis, requiring manual transcription of every record. We investigate whether state-of-the-art document analysis models can address this task, and introduce CENTURIA, a dataset of 507 pottery records from the Roman site of Carnuntum, providing transcriptions, bounding boxes, and structured field-level labels across seven metadata categories. Benchmarking five OCR models reveals a substantial domain gap: zero-shot transcription error reaches 15-32% SpACER-M, far exceeding rates on printed archival documents, with domain-specific fields recovered in fewer than 3% of cases. LoRA fine-tuning on just 57 samples, reflecting a realistic archival annotation budget, closes this gap, reducing transcription error to below 1.5% and recovering overall field-level accuracy above 87%. Our results show that a small expert-validated fine-tuning set suffices to convert handwritten pottery documentation into structured, searchable metadata ready for archaeological databases.

OCR考古数据手写识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。