arXiv:2606.01498cs.CLcs.AI2026-06被引 2

构建多轮对话时间序列推理评测集,检验大模型持续分析能力

TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

论文配图:TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
图 1 · 摘自论文原文
  • 将真实数据转为带可验证答案的多轮对话,模拟复杂分析流程
  • 240个任务中决策类任务性能下降超60%,暴露记忆与不确定性处理缺陷
  • 适合评估大模型在金融、医疗等领域的长期推理能力

时间序列数据支撑众多现实场景的关键决策。尽管大语言模型代理可通过自然语言和工具分析数据,但其在多轮对话中进行可靠时间序列分析的能力尚不明确。现有基准仅关注单步任务如预测和异常检测,忽略了用户目标演变、代理需基于前期分析、结论随证据积累形成的实际工作流。本文提出TimeSage-MT,一个包含240个任务、2,680轮对话的多轮代理时间序列推理基准,覆盖8个真实世界领域,从基础探索到决策导向分析。该基准通过可复现流程将真实时间序列数据转化为带可验证答案的多轮对话,提供统一评估协议与公开排行榜。为验证其有效性,我们评估了前沿大模型及新提出的结构化代理TimeSage(配备完整时间序列技能库)。结果表明,在决策类任务上性能显著下降,主因是记忆缺失、不确定性处理不足与领域决策偏差。TimeSage-MT揭示了当前代理推理中的关键差距,并为后续发展提供了严谨基础。

原文摘要 · Abstract (English)

Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural language and tools, it remains unclear whether they can conduct reliable time series analysis across multi-turn conversations. Existing benchmarks focus on single-step tasks such as forecasting and anomaly detection, overlooking practical workflows where user goals evolve, agents must build on prior analyses, and conclusions emerge from accumulated evidence. In this work, we introduce TimeSage-MT, a multi-turn benchmark for agentic time series reasoning with 240 tasks and 2,680 dialogue turns across 8 real-world domains, spanning basic exploration to decision-oriented analysis. TimeSage-MT is built through a reproducible pipeline that converts real-world time series data into multi-turn conversations with verifiable answers. It provides a unified evaluation protocol and public leaderboard for comparing time series agentic systems. To demonstrate the benchmark's utility, we evaluate frontier LLMs alongside TimeSage, a novel structured agent equipped with a comprehensive time series skill library. The results show sharp performance drops on decision-oriented tasks, driven by failures in memory, uncertainty handling, and domain-based decision making. TimeSage-MT exposes critical gaps in current agentic reasoning and provides a rigorous foundation for future development.

时间序列大模型代理多轮推理评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。