构建首个针对出土古籍的综合性评估基准,填补古文字理解评测空白。
AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora
- 从字形、读音、词义到语境四维度设计十项任务,覆盖出土古籍理解全链条
- 在10项任务上测试主流大模型,发现其在古籍理解上潜力巨大但远未超越人类
- 联合考古专家验证,为古文字研究与AI应用提供可复用评估框架
古籍理解对考古学和中华文明研究至关重要。随着大语言模型快速发展,亟需能评估其古文字理解能力的基准。现有中文基准多聚焦现代汉语及传世文献,缺乏对出土文献的覆盖。为此,我们提出 AncientBench,旨在评估大模型对古文字的理解能力,尤其关注出土文献场景。该基准涵盖字形、读音、词义和语境四个维度,包含十项任务,如部首识别、声旁判断、同音字辨析、填空、翻译等,形成系统性评估框架。我们召集考古学者进行实验评估,提出一个古文基础模型,并在当前性能最优的大语言模型上开展广泛实验。结果表明,大模型在古籍理解中展现出显著潜力,但仍与人类存在明显差距。本研究致力于推动大模型在考古学与古汉语领域的应用与发展。
原文摘要 · Abstract (English)
Comprehension of ancient texts plays an important role in archaeology and understanding of Chinese history and civilization. The rapid development of large language models needs benchmarks that can evaluate their comprehension of ancient characters. Existing Chinese benchmarks are mostly targeted at modern Chinese and transmitted documents in ancient Chinese, but the part of excavated documents in ancient Chinese is not covered. To meet this need, we propose the AncientBench, which aims to evaluate the comprehension of ancient characters, especially in the scenario of excavated documents. The AncientBench is divided into four dimensions, which correspond to the four competencies of ancient character comprehension: glyph comprehension, pronunciation comprehension, meaning comprehension, and contextual comprehension. The benchmark also contains ten tasks, including radical, phonetic radical, homophone, cloze, translation, and more, providing a comprehensive framework for evaluation. We convened archaeological researchers to conduct experimental evaluations, proposed an ancient model as baseline, and conducted extensive experiments on the currently best-performing large language models. The experimental results reveal the great potential of large language models in ancient textual scenarios as well as the gap with humans. Our research aims to promote the development and application of large language models in the field of archaeology and ancient Chinese language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。