arXiv:2601.09523cs.IR2026-01被引 8

首个融合时间推理与复杂检索的多领域基准,挑战模型跨时段证据整合能力。

TEMPO: A Realistic Multi-Domain Benchmark for Temporal Reasoning-Intensive Retrieval

  • 构建13个领域、1730个需深度时间推理的复杂查询
  • 设计3976步分解检索流程,支持多跳评估
  • 引入时间覆盖度等新指标,评估结果时间跨度完整性

现有时间问答基准多聚焦新闻语料中的简单事实查询,而推理密集型检索基准缺乏时间定位。然而真实信息需求常需对时间演进进行推理,并跨时间段整合证据。我们提出TEMPO,首个结合时间推理与推理密集型检索的多领域基准,涵盖13个领域。TEMPO包含:(1)1,730个复杂查询,要求深度时间推理,如追踪变化、识别趋势或跨时期证据对比;(2)3,976个分步检索规划,每步对应黄金文档,支持多跳评估;(3)新颖的时间度量指标,包括Temporal Coverage@k和Temporal Precision@k,衡量结果是否覆盖所需时间范围。对12个检索系统的评估显示巨大挑战:最佳模型DiVeR仅达32.0 NDCG@10和71.4% Temporal Coverage@10,表明获取完整时间证据仍极困难。我们认为TEMPO为提升检索与RAG系统中的时间推理能力提供了严峻基准。代码与数据已开源:https://github.com/tempo-bench/Tempo,官网:https://tempo-bench.github.io/

原文摘要 · Abstract (English)

Existing temporal QA benchmarks focus on simple fact-seeking queries from news corpora, while reasoning-intensive retrieval benchmarks lack temporal grounding. However, real-world information needs often require reasoning about temporal evolution and synthesizing evidence across time periods. We introduce TEMPO, the first benchmark combining temporal reasoning with reasoning-intensive retrieval across 13 domains. TEMPO features: (1) 1,730 complex queries requiring deep temporal reasoning such as tracking changes, identifying trends, or comparing cross-period evidence; (2) step-wise retrieval planning with 3,976 decomposed steps and gold documents mapped to each step for multi-hop evaluation; and (3) novel temporal metrics including Temporal Coverage@k and Temporal Precision@k measuring whether results span required time periods. Evaluation of 12 retrieval systems reveals substantial challenges: the best model (DiVeR) achieves only 32.0 NDCG@10 and 71.4\% Temporal Coverage@10, demonstrating difficulty in retrieving temporally complete evidence. We believe TEMPO provides a challenging benchmark for improving temporal reasoning in retrieval and RAG systems. Our code and data are available at https://github.com/tempo-bench/Tempo. See also our official website: https://tempo-bench.github.io/.

时间推理检索评估多跳检索基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。