arXiv:2608.27169cs.CV2026-08中稿 · EMNLP

构建涵盖三千年古籍的多介质多字体识别基准,解决古文字识别评估碎片化问题

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

论文配图:Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition
图 1 · 摘自论文原文
  • 覆盖3000年历史、9类文物、7种字体,构建多维度古文识别数据集
  • 在2700张图像上测试发现主流模型仍存在字形变异与幻觉等根本性难题
  • 针对不同材质设计标注标准,支持公平跨介质评估,适合文化遗产数字化研究者

古代中文文物文本识别对文化遗产数字化至关重要,现有基准因时间覆盖有限、介质类型单一、字体不全而存在碎片化问题。为此,我们提出 Ancient-Bench,一个包含2700张图像的综合性基准,涵盖三个维度:跨三千年的字符演变、九类文物介质、七种历史书写体。为实现异质介质间一致且公平的评估,我们针对不同介质特性定义了三种标注标准:符号标准化、字符标准化与解析标准化。在该基准上对通用视觉语言模型(VLMs)和专用OCR模型的广泛实验表明,古中文文物文本识别仍未根本解决,持续面临变体字符、专有符号及生成幻觉等挑战。数据集已开源:https://github.com/SCUT-DLVCLab/Ancient_Bench。

原文摘要 · Abstract (English)

Ancient Chinese artifact text recognition is fundamental to heritage digitization, and benchmarks for ancient texts are essential for evaluating current model capabilities. However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types. Therefore, we present Ancient-Bench, a comprehensive benchmark of 2,700 images for ancient Chinese artifact text recognition, featuring three dimensions: Multi-millennial (spanning 3,000 years of character evolution), Multi-medium (covering nine artifact categories), and Multi-script (encompassing seven historical script forms). To enable consistent and fair evaluation across heterogeneous media, we further define three annotation standards tailored to the medium-specific characteristics of ancient texts: symbol standardization, character standardization, and parsing standardization. Extensive experiments on Ancient-Bench covering general Vision-Language Models (VLMs) and OCR-specialist models reveal that ancient Chinese artifact text recognition remains fundamentally unsolved, with persistent challenges in variant characters, specialized symbols, and hallucination. The dataset is available at https://github.com/SCUT-DLVCLab/Ancient_Bench.

古文字识别文化遗产多模态数据集OCR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。