构建首个针对不规则时间序列的智能体问答基准,测试大模型真实场景表现。
Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning

- 设计可独立使用的不规则时间序列问答基准,支持标准化评估
- 包含1700个问题,覆盖13个领域和10类任务,反映真实数据复杂性
- 适合研究大模型在真实时间序列分析中的推理与工具使用能力
现实应用中的时间序列数据普遍呈现不规则特征:观测异步、缺失值具有信息量而非随机,不同传感器和运行时段采样频率各异。然而现有时间序列问答(TSQA)基准大多假设规则采样输入,导致对大语言模型(LLMs)和智能体在不规则条件下的表现缺乏理解。为填补这一空白,我们提出IRTS-ToolBench,一个包含1700个问题、覆盖13个领域和10类任务的基准。该基准可由任何从事基于大模型的不规则时间序列分析的研究者独立使用,提供标准化输入和可复现的评估协议。代码已开源:https://github.com/SanhornC/IRTS-ToolBench。
原文摘要 · Abstract (English)
Time series data in real-world deployments is overwhelmingly irregular. Observations are asynchronous, missing values are informative rather than random, and sampling frequencies vary across sensors and operational windows. However, existing Time Series Question Answering (TSQA) benchmarks mostly assume regularly sampled inputs, leaving a fundamental gap in understanding how large language models (LLMs) and AI agents perform under irregular conditions. To bridge this gap, we introduce IRTS-ToolBench, a benchmark of 1,700 questions spanning 10 task types across 13 domains. IRTS-ToolBench is designed to be used independently by any researcher working on LLM-based irregular time series analysis, providing standardized inputs and a reproducible evaluation protocol. Code can be found in https://github.com/SanhornC/IRTS-ToolBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。