arXiv:2509.04432cs.CL2025-09中稿 · IJCNLP-AACL 2025被引 2

测试大模型对日本和历的处理能力,发现主流模型表现不佳。

Can Language Models Handle a Non-Gregorian Calendar? The Case of the Japanese wareki

  • 构建含和历日期的时序推理数据集进行评估
  • GPT-4o等模型在和历计算上错误率超30%
  • 适合关注文化特定时序理解的开发者参考

时间推理与知识是语言模型的重要能力。尽管已有大量研究关注语言模型在公历下的时间推理表现,但对非公历系统(如日本和历、伊斯兰历、希伯来历)的处理仍缺乏系统评估。本文针对日本和历开展系统性评估,构建需运用和历日期进行时序推理的数据集。测试开放与闭源模型后发现,部分模型可完成历法转换,但包括GPT-4o、Deepseek V3及日语专用模型在内的多个主流模型在和历算术与相关知识理解上表现不佳。误差分析表明,和历表达在训练语料中频率低以及模型知识中的公历偏差可能是原因。结果凸显了提升语言模型在文化特定任务(如历法理解)方面能力的重要性。

原文摘要 · Abstract (English)

Temporal reasoning and knowledge are essential capabilities for language models (LMs). While much prior work has analyzed and improved temporal reasoning in LMs, most studies have focused solely on the Gregorian calendar. However, many non-Gregorian systems, such as the Japanese, Hijri, and Hebrew calendars, are in active use and reflect culturally grounded conceptions of time. If and how well current LMs can accurately handle such non-Gregorian calendars has not been evaluated so far. Here, we present a systematic evaluation of how well language models handle one such non-Gregorian system: the Japanese wareki. We create datasets that require temporal knowledge and reasoning in using wareki dates. Evaluating open and closed LMs, we find that some models can perform calendar conversions, but GPT-4o, Deepseek V3, and even Japanese-centric models struggle with Japanese calendar arithmetic and knowledge involving wareki dates. Error analysis suggests corpus frequency of Japanese calendar expressions and a Gregorian bias in the model's knowledge as possible explanations. Our results show the importance of developing LMs that are better equipped for culture-specific tasks such as calendar understanding.

时序推理和历多文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。