测试大模型在真实时间任务中的上下文理解能力,发现预测准确不等于真懂时间逻辑。
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
- 分四层评估模型对历史数据、上下文和事件的综合理解能力
- 现有模型在上下文推理中表现差,数字准确但逻辑断裂
- 适合研究时序建模、智能体决策或需理解事件影响的研究者
当前强预测性能是否反映真正的时序理解,还是仅依赖上下文与事件驱动条件尚不明确。我们提出TemporalBench,一个跨零售、医疗、能源和物理系统的多领域基准,用于评估大模型在逐步丰富信息设置下的时序推理行为。该基准采用四层任务分类:历史结构解读、无上下文预测、上下文时序推理、事件条件预测。通过控制未来目标与上下文信息的可访问性,实现对模型能否正确识别时序模式、与外部上下文对齐、在条件变化时调整预测的诊断分析。大量基线实验表明,高数值预测准确率并不能可靠转化为稳健的上下文或事件感知时序推理能力;现有智能体框架存在能力碎片化与系统性失效模式,这些缺陷在仅关注预测的基准下难以察觉。TemporalBench数据集已公开于https://huggingface.co/datasets/Melady/TemporalBench,同时提供公共排行榜https://huggingface.co/spaces/Melady/TemporalBench_Leaderboard。
原文摘要 · Abstract (English)
It is unclear whether strong forecasting performance reflects genuine temporal understanding or the ability to reason under contextual and event-driven conditions. We introduce TemporalBench, a multi-domain benchmark designed to evaluate temporal reasoning behavior under progressively richer informational settings. TemporalBench adopts a four-tier task taxonomy that examines historical structure interpretation, context-free forecasting, contextual temporal reasoning, and event-conditioned prediction across four real-world domains: retail, healthcare, energy, and physical systems. By controlling access to future targets and contextual information, the benchmark enables a diagnostic analysis of whether models can correctly interpret temporal patterns, align them with external context, and adapt predictions when conditions change. Extensive baseline experiments show that strong numerical forecasting accuracy does not reliably translate into robust contextual or event-aware temporal reasoning; instead, existing agent frameworks exhibit fragmented strengths and systematic failure modes that remain largely hidden under forecasting-only benchmarks. The TemporalBench dataset is publicly available at https://huggingface.co/datasets/Melady/TemporalBench, and we additionally provide a public leaderboard at https://huggingface.co/spaces/Melady/TemporalBench_Leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。