arXiv:2506.03984cs.CL2025-06ACL被引 7

首次系统评估大模型在时空联合推理能力,发现地理信息连接仍存短板。

Around the World in 24 Hours: Probing LLM Knowledge of Time and Place

  • 构建涵盖289城37时区的GeoTemp数据集,测试模型时空联合推理能力。
  • 模型在纯时间推理上表现良好,但跨时空关联任务性能显著下降。
  • 低困惑度地名更易被正确识别,提示训练数据重复性影响表现。

时空推理是理解世界的核心能力,但现有研究多孤立考察时间或空间推理,且局限于简单或人工环境。本文首次系统评估大语言模型在时空联合推理方面的能力。为此,我们构建了包含320,000个提示的GeoTemp数据集,覆盖217个国家的289个城市及37个时区。基于此,我们评估了三种模型家族的八款开源聊天模型在不同时间与地理知识组合下的表现。结果表明,多数模型在仅涉及时间推理的任务中表现良好,且性能随模型规模提升而改善;但在需同时整合时空信息的任务中,性能受限。我们未发现特定地理区域对性能有明显影响,但发现低困惑度(low perplexity)的地名表现更好,暗示其在训练中出现频率更高。此外,提示设计影响显著:直接注入地理信息可提升表现,而链式思维提示反而在简单任务中降低效果。

原文摘要 · Abstract (English)

Reasoning over time and space is essential for understanding our world. However, the abilities of language models in this area are largely unexplored as previous work has tested their abilities for logical reasoning in terms of time and space in isolation or only in simple or artificial environments. In this paper, we present the first evaluation of the ability of language models to jointly reason over time and space. To enable our analysis, we create GeoTemp, a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones. Using GeoTemp, we evaluate eight open chat models of three different model families for different combinations of temporal and geographic knowledge. We find that most models perform well on reasoning tasks involving only temporal knowledge and that overall performance improves with scale. However, performance remains constrained in tasks that require connecting temporal and geographical information. We do not find clear correlations of performance with specific geographic regions. Instead, we find a significant performance increase for location names with low model perplexity, suggesting their repeated occurrence during model training. We further demonstrate that their performance is heavily influenced by prompt formulation - a direct injection of geographical knowledge leads to performance gains, whereas, surprisingly, techniques like chain-of-thought prompting decrease performance on simpler tasks.

时空推理大模型评估地理知识提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。