专为历史推理设计的基准与智能体,显著提升AI对多模态历史数据的理解能力。
On Path to Multimodal Historical Reasoning: HistBench and HistAgent
- 构建覆盖29种语言、414个问题的历史推理基准HistBench
- 历史专用智能体HistAgent在GPT-4o上达27.54%准确率,超越通用模型
- 适合历史研究、跨语言分析及多模态推理方向的研究者使用
大型语言模型(LLMs)在多个领域取得显著进展,但在人文学科尤其是历史领域的应用仍不充分。历史推理涉及多模态源理解、时间推断和跨语言分析,现有通用智能体缺乏领域专业知识。为此,我们推出HistBench,一个由40多位专家编写、包含414个高质量问题的新基准,涵盖从原始文献事实检索到手稿与图像解释,再到考古学、语言学等跨学科挑战。数据集覆盖29种古今语言,横跨多个历史时期与地区。在该基准上,主流模型表现不佳。因此,我们提出HistAgent——一个专为历史设计的智能体,配备精心设计的OCR、翻译、档案检索与图像理解工具。基于GPT-4o的HistAgent在HistBench上达到27.54% pass@1与36.47% pass@2,显著优于其他模型,包括GPT-4o(18.60%)、DeepSeek-R1(14.49%)和Open Deep Research-smolagents(20.29% pass@1,25.12% pass@2),凸显通用模型局限性,并验证了领域专用智能体的优势。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents perform well on many existing benchmarks, they lack the domain-specific expertise required to engage with historical materials and questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality questions designed to evaluate AI's capacity for historical reasoning and authored by more than 40 expert contributors. The tasks span a wide range of historical problems-from factual retrieval based on primary sources to interpretive analysis of manuscripts and images, to interdisciplinary challenges involving archaeology, linguistics, or cultural history. Furthermore, the benchmark dataset spans 29 ancient and modern languages and covers a wide range of historical periods and world regions. Finding the poor performance of LLMs and other agents on HistBench, we further present HistAgent, a history-specific agent equipped with carefully designed tools for OCR, translation, archival search, and image understanding in History. On HistBench, HistAgent based on GPT-4o achieves an accuracy of 27.54% pass@1 and 36.47% pass@2, significantly outperforming LLMs with online search and generalist agents, including GPT-4o (18.60%), DeepSeek-R1(14.49%) and Open Deep Research-smolagents(20.29% pass@1 and 25.12% pass@2). These results highlight the limitations of existing LLMs and generalist agents and demonstrate the advantages of HistAgent for historical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。