用科举制度测试大模型历史推理能力,发现顶尖模型仍有明显短板。
Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

- 基于中国科举制度构建专业历史推理评测基准
- 18个大模型在400道题上平均得分不足及格线
- 适合研究历史推理与AI评估的学者使用
尽管大型语言模型在文本处理等历史任务中应用日益广泛,但其在专业级历史推理方面的能力仍缺乏系统评估。现有基准多聚焦于知识广度或词汇理解,未能涵盖历史研究中的核心高阶能力,如证据推理。为此,我们提出ProHist-Bench,一个以持续1300余年的中国科举制度为依托的新型评测基准,全面反映东亚政治、社会与思想史。该基准由跨学科专家共同设计,包含覆盖八朝的400道高难度题目,配套10,891条细粒度评分标准。对18个LLM的严格评测显示,即使最先进模型在复杂历史研究问题上仍表现不佳,存在显著能力差距。我们期望ProHist-Bench能推动领域专用推理型语言模型的发展,促进计算历史学进步,并进一步挖掘大模型的潜力。数据集已开源:https://github.com/inclusionAI/ABench/tree/main/ProHist-Bench。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level historical reasoning remains underexplored. Existing benchmarks primarily assess basic knowledge breadth or lexical understanding, failing to capture the higher-order skills, such as evidentiary reasoning,that are central to historical research. To fill this gap, we introduce ProHist-Bench, a novel benchmark anchored in the Chinese Imperial Examination (Keju) system, a comprehensive microcosm of East Asian political, social, and intellectual history spanning over 1,300 years. Developed through deep interdisciplinary collaboration, ProHist-Bench features 400 challenging, expert-curated questions across eight dynasties, accompanied by 10,891 fine-grained evaluation rubrics. Through a rigorous evaluation of 18 LLMs, we reveal a significant proficiency gap: even state-of-the-art LLMs struggle with complex historical research questions. We hope ProHist-Bench will facilitate the development of domain-specific reasoning LLMs, advance computational historical research, and further uncover the untapped potential of LLMs. We release ProHist-Bench at https://github.com/inclusionAI/ABench/tree/main/ProHist-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。