测试大模型在波兰历史高考中的真实表现,发现其历史推理存在明显短板。
Lost in Historical Time? A Polish History Matura Benchmark for Large Language Models

- 用波兰高中毕业考题评估8个主流大模型的历史理解能力
- 模型总分接近满分,但不同题型表现差异大,波兰史题得分明显偏低
- 揭示模型常犯错:误读史料和时间错位,适合教育研究者参考
大语言模型被学生广泛用作知识来源,但现有评测很少考察其历史推理能力。本文评估了八个主流LLM在波兰高中毕业考试(Matura)历史科目的表现,涵盖2023–2025年三套官方试卷,包含简答题与论述题,并与真实考生群体对比。尽管模型整体得分接近满分,但聚合分数掩盖了其在任务类型、信息模态和地理范围上的不稳定性,对波兰本土历史内容始终存在系统性扣分。定性分析发现两类典型错误:一是将史料内容当作事实而非分析对象(源脱嵌),二是回答出现时间错位(时间迷失)。本研究首次构建基于波兰国家课程的历史类大模型评测基准。
原文摘要 · Abstract (English)
Language models are widely used by students as knowledge sources, yet benchmarks rarely assess their interpretative historical reasoning. We evaluate eight leading LLMs on the Polish high school exit exam (Matura) in history - three official papers from 2023-2025, comprising short-answer questions and extended essays - and compare model performance against the human examinee population. Although models score near the ceiling, aggregate scores mask distinct competency profiles: rankings are unstable across task types, source modalities, and geographical scopes, with a consistent penalty for Polish versus Global history content. Qualitative error analysis reveals two recurring failure modes - source decontextualization, when models reason from source content rather than treating it as an object of analysis, and temporal disorientation, when responses are historically misplaced. This study introduces the first LLM history benchmark grounded in the Polish national curriculum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。