首个评估汉字跨时代视觉感知能力的基准,揭示大模型在古文字识别中的短板。
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

- 构建跨千年汉字演变的标准化图像数据集,覆盖七种主要书体
- 包含2800张平衡样本,涵盖甲骨、简牍到纸本书法等多样载体
- 提出自适应标注范式,专为应对历史字形剧烈变化设计
视觉大语言模型(VLLMs)在现代文本密集视觉理解中表现卓越,但在面对汉字书写系统持续形态演变时的感知鲁棒性仍缺乏研究。现有古文字数据集多聚焦孤立历史时期,未能捕捉跨越数千年的系统性视觉分布变化。为填补这一空白并推动数字人文发展,我们提出Chronicles-OCR,首个专门评估VLLMs在汉字完整演化轨迹(即七种中国书体)上跨时代视觉感知能力的综合性基准。该数据集由顶尖机构专家协作构建,包含2800张严格平衡的图像,涵盖龟甲、简牍至纸本书法等多种物理媒介。针对不同历史阶段剧烈的形态与拓扑变化,提出一种新型阶段自适应标注范式。基于此,Chronicles-OCR 设定四项严谨量化任务:跨时期字符定位、基于视觉指代的细粒度古字识别、古文解析及书体分类。通过分离视觉感知与语义推理,该基准为暴露当前VLLMs在历史文本感知中的局限性提供权威平台,推动具备演化意识的鲁棒性历史文字理解发展。数据集已公开于 https://github.com/VirtualLUOUCAS/Chronicles-OCR。
原文摘要 · Abstract (English)
Vision Large Language Models (VLLMs) have achieved remarkable success in modern text-rich visual understanding. However, their perceptual robustness in the face of the continuous morphological evolution of historical writing systems remains largely unexplored. Existing ancient text datasets typically focus on isolated historical periods, failing to capture the systematic visual distribution shifts spanning thousands of years. To bridge this gap and empower Digital Humanities, we introduce Chronicles-OCR, the first comprehensive benchmark specifically designed to evaluate the cross-temporal visual perception capabilities of VLLMs across the complete evolutionary trajectory of Chinese characters, known as the Seven Chinese Scripts. Curated in collaboration with top-tier institutional domain experts, the dataset comprises 2,800 strictly balanced images encompassing highly diverse physical media, ranging from tortoise shells to paper-based calligraphy. To accommodate the drastic morphological and topological variations across different historical stages, we propose a novel Stage-Adaptive Annotation Paradigm. Based on this, Chronicles-OCR formulates four rigorous quantitative tasks: cross-period character spotting, fine-grained archaic character recognition via visual referring, ancient text parsing, and script classification. By isolating visual perception from semantic reasoning, Chronicles-OCR provides an authoritative platform to expose the limitations of current VLLMs, paving the way for robust, evolution-aware historical text perception. Chronicles-OCR is publicly available at https://github.com/VirtualLUOUCAS/Chronicles-OCR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。