arXiv:2504.07646cs.CLcs.AI2025-04被引 2

测试大模型在匿名数据上的时间推理能力,发现仅靠大模型不够可靠。

On the Temporal Question-Answering Capabilities of Large Language Models Over Anonymized Data

  • 构建RATA数据集,用匿名半结构化数据评估时间推理。
  • 对比多种方法,发现现有技术在时间推理上表现有限。
  • 强调需整合多技术才能实现稳定可靠的时间推理。

大型语言模型(LLMs)在训练数据之外的时间推理任务中的适用性仍待探索。本文聚焦于结构化和半结构化匿名数据,不仅设计了直接的LLM处理流程,还比较了多种方法并进行了深入分析。我们识别并研究了自然语言中17种常见的时间推理任务及其算法组件。为评估模型性能,创建了名为《推理与时间回答能力》(RATA)的数据集,采用半结构化匿名数据以确保依赖推理而非先验知识。我们对比了包括思维树(Tree-of-Thought)、自我反思(self-reflection)和代码执行在内的多种前沿技术,并针对此场景进行了调优。结果表明,实现可扩展且可靠的解决方案,仅靠独立的LLM是不够的,亟需集成式方法。

原文摘要 · Abstract (English)

The applicability of Large Language Models (LLMs) in temporal reasoning tasks over data that is not present during training is still a field that remains to be explored. In this paper we work on this topic, focusing on structured and semi-structured anonymized data. We not only develop a direct LLM pipeline, but also compare various methodologies and conduct an in-depth analysis. We identified and examined seventeen common temporal reasoning tasks in natural language, focusing on their algorithmic components. To assess LLM performance, we created the \textit{Reasoning and Answering Temporal Ability} dataset (RATA), featuring semi-structured anonymized data to ensure reliance on reasoning rather than on prior knowledge. We compared several methodologies, involving SoTA techniques such as Tree-of-Thought, self-reflexion and code execution, tuned specifically for this scenario. Our results suggest that achieving scalable and reliable solutions requires more than just standalone LLMs, highlighting the need for integrated approaches.

时间推理大模型数据匿名

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。