首个中文古籍视觉语言评测基准,覆盖3000页古籍多任务评估。
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
- 构建从文字识别到知识推理的五项任务体系
- 涵盖14类古籍、超100部书、约3000页真实文献
- 适配研究古籍数字化与跨模态理解的学者
中文古籍是承载千年历史文化的重要载体,蕴含丰富知识但面临数字化与理解难题:传统方法仅扫描图像,而现有视觉语言模型(VLMs)难以应对古籍复杂的视觉与语言特征。现有文档评测基准集中于英文印刷体或简体中文,缺乏对古汉语文献的评估体系。为此,我们提出AncientDoc,首个面向中文古籍的评测基准,旨在评估VLMs从光学字符识别(OCR)到知识推理的全链路能力。AncientDoc包含五项任务:页面级OCR、白话翻译、基于推理的问答、基于知识的问答、语言变体问答,覆盖14类文献、超过100部典籍,约3000页文本。基于该基准,我们采用多指标评估主流VLMs,并借助人工对齐的大语言模型进行评分。
原文摘要 · Abstract (English)
Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while current Vision-Language Models (VLMs) struggle with their visual and linguistic complexity. Existing document benchmarks focus on English printed texts or simplified Chinese, leaving a gap for evaluating VLMs on ancient Chinese documents. To address this, we present AncientDoc, the first benchmark for Chinese ancient documents, designed to assess VLMs from OCR to knowledge reasoning. AncientDoc includes five tasks (page-level OCR, vernacular translation, reasoning-based QA, knowledge-based QA, linguistic variant QA) and covers 14 document types, over 100 books, and about 3,000 pages. Based on AncientDoc, we evaluate mainstream VLMs using multiple metrics, supplemented by a human-aligned large language model for scoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。