arXiv:2511.04432cs.CL2025-11被引 1

让大模型模拟1940年视角回答历史问题,测试其时间推理能力。

If I Could Turn Back Time: Temporal Reframing as a Historical Reasoning Task for LLMs

  • 用1940年挪威书籍的题目,让大模型以当时视角作答。
  • 英语提问比挪威语效果更好,大模型规模越大表现越优。
  • 适合研究历史理解与多语言大模型性能的学者参考。

本研究探究大语言模型进行时间推理的能力。基于一本1940年出版的挪威通俗读物中的趣味问题,我们引导模型以1940年视角作答,并在英语和挪威语双语环境下测试。正确答案常为完整句子,评分采用大模型作为评判者,辅以母语者抽样验证。结果显示,英语提示下的表现优于挪威语提示,这一结果出乎意料;同时,使用更大规模的模型显著提升准确率。实验涵盖DeepSeek-R1、Gemma3、Qwen3、Llama3.1等模型家族,以及专为挪威语设计的最大可用大模型。

原文摘要 · Abstract (English)

In this study, we experiment with the ability of LLMs to do temporal reasoning. Using a Norwegian book from 1940 containing trivia questions, we prompt the LLMs to answer the questions as if it were 1940. We also pose the questions in both English and Norwegian. Correct answers are often presented as sentences, and grading is done by means of LLM-as-judge, with sampled checks by a native speaker. Prompting in English consistently gave better results than in Norwegian, an unexpected result. In contrast, using larger LLMs improved results. We tested the DeepSeek-R1, Gemma3, Qwen3, and Llama3.1 model families, and also the largest available LLM especially crafted for Norwegian.

时间推理多语言历史理解大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。