arXiv:2605.24703cs.CLcs.AI2026-05

构建可分解的时间序列分析能力评测基准,揭示模型在时序推理上的真实短板。

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

论文配图:TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering
图 1 · 摘自论文原文
  • 按三种可组合技能设计可控评测任务:时序尺度选择、时间定位、跨区间整合。
  • 十款主流大模型在跨区间整合技能上表现普遍薄弱,非智能体模型尤其明显。
  • 适合研究时序推理、模型评估或想提升时间序列问答能力的开发者与研究人员。

大型语言模型(LLMs)和时序语言模型(TSLMs)正被广泛应用于时间序列问答(TSQA)。与纯文本问答不同,TSQA要求模型基于具有多尺度、特定时间位置或分离间隔特征的时序信号进行推理。然而,现有基准多按任务类型或高层次推理类别组织,难以诊断驱动性能的底层信号级能力。本文提出TS-Skill,一个用于评估三种可组合分析技能的受控基准:时序尺度选择(SK1)、时间定位(SK2)和跨区间整合(SK3)。该基准包含带时间戳的问题、广泛的领域覆盖以及人工验证的问答质量。为规模化构建,我们开发了SKEvol——一种技能引导的智能体框架,结合领域感知的时序种子生成、技能控制的问题生成、元数据与代码辅助的答案构建、多阶段信号锚定验证及人机协同校准。对十种前沿LLMs和TSLMs的实验显示,各模型在SK1-SK3上存在显著且不均衡的能力差距。尤其在跨区间整合(SK3)方面,非智能体模型始终表现不佳,而工具增强型智能体在独立处理SK3时展现出选择性优势。结果表明,技能级评估能揭示被整体评分掩盖的时序推理缺陷。

原文摘要 · Abstract (English)

Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-only QA, TSQA requires models to ground answers in temporal signals whose patterns may occur at different scales, specific time locations, or across separated intervals. However, existing benchmarks are typically organized by task types or high-level reasoning categories, making it difficult to diagnose the underlying signal-level capabilities driving model performance. We introduce TS-Skill, a controlled benchmark for evaluating three composable analytical skills in TSQA: temporal scale selection (SK1), temporal localization (SK2), and cross-interval integration (SK3). TS-Skill provides timestamp-aware questions, broad domain coverage, and human-validated QA quality. To construct the benchmark at scale, we develop SKEvol, a skill-guided agentic framework that combines domain-aware time-series seed generation, skill-controlled question generation, metadata- and code-assisted answer construction, multi-phase signal-grounded verification, and human-in-the-loop curation. Experiments on ten state-of-the-art LLMs and TSLMs reveal substantial and uneven capability gaps across SK1-SK3. In particular, SK3 remains consistently challenging for non-agent models, whereas tool-augmented agents show a selective advantage on standalone SK3. These findings demonstrate that skill-level evaluation can uncover temporal reasoning failures that are obscured by aggregate TSQA scores.

时序推理模型评估问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。