构建时序表格数据集,评估大模型对动态事件的推理能力
TransientTables: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
- 用模板生成法构建3971个时序推理问题
- 在14000张跨时段表格上验证模型表现
- 通过任务分解提升大模型时序推理效果
人类持续产生新发现,理解事件发展的时间序列对推动科学与社会发展至关重要。然而,大语言模型通常基于静态数据训练,难以有效进行时序推理。为此,我们构建了TRANSIENTTABLES数据集,包含3971个问题,源自超过14,000张表格,覆盖1238个实体在多个时间周期内的演变。我们提出基于模板的问题生成流程,利用大模型迭代优化模板与问题。同时,采用先进大模型建立基线,并引入以任务分解为核心的新型建模策略,显著提升模型性能。
原文摘要 · Abstract (English)
Humans continuously make new discoveries, and understanding temporal sequence of events leading to these breakthroughs is essential for advancing science and society. This ability to reason over time allows us to identify future steps and understand the effects of financial and political decisions on our lives. However, large language models (LLMs) are typically trained on static datasets, limiting their ability to perform effective temporal reasoning. To assess the temporal reasoning capabilities of LLMs, we present the TRANSIENTTABLES dataset, which comprises 3,971 questions derived from over 14,000 tables, spanning 1,238 entities across multiple time periods. We introduce a template-based question-generation pipeline that harnesses LLMs to refine both templates and questions. Additionally, we establish baseline results using state-of-the-art LLMs to create a benchmark. We also introduce novel modeling strategies centered around task decomposition, enhancing LLM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。