构建首个覆盖6类时序分析任务的统一评测基准
TSAQA: Time Series Analysis Question And Answering Benchmark
- 设计包含真伪、多选、谜题三类格式的统一评测框架
- 涵盖210k样本,跨13个领域,零样本下最优模型仅65.08分
- 适合评估大模型在复杂时序理解上的能力,尤其关注中文场景
时间序列数据在金融、医疗、交通和环境科学等关键领域中至关重要。尽管近期已有研究探索多任务时间序列问答(QA),但现有基准仍局限于预测与异常检测任务。我们提出TSAQA,一个新型统一基准,旨在拓展任务覆盖面并评估多样化的时序分析能力。TSAQA在单一框架下整合了六种不同任务,涵盖常规分析(如异常检测、分类)和高级分析(如特征刻画、对比、数据转换、时序关系分析)。数据集覆盖210k样本,涉及13个领域,采用多种格式:真伪判断(TF)、多选题(MC)及一种新颖的谜题(PZ)形式,全面评估时序分析能力。零样本评估显示,当前大型语言模型(LLMs)面临挑战:表现最佳的商用模型Gemini-2.5-Flash平均得分仅为65.08。尽管指令微调提升了开源模型性能,但最佳开源模型LLaMA-3.1-8B仍有显著提升空间,凸显时序分析对大模型的复杂性。
原文摘要 · Abstract (English)
Time series data are integral to critical applications across domains such as finance, healthcare, transportation, and environmental science. While recent work has begun to explore multi-task time series question answering (QA), current benchmarks remain limited to forecasting and anomaly detection tasks. We introduce TSAQA, a novel unified benchmark designed to broaden task coverage and evaluate diverse temporal analysis capabilities. TSAQA integrates six diverse tasks under a single framework ranging from conventional analysis, including anomaly detection and classification, to advanced analysis, such as characterization, comparison, data transformation, and temporal relationship analysis. Spanning 210k samples across 13 domains, the dataset employs diverse formats, including true-or-false (TF), multiple-choice (MC), and a novel puzzling (PZ), to comprehensively assess time series analysis. Zero-shot evaluation demonstrates that these tasks are challenging for current Large Language Models (LLMs): the best-performing commercial LLM, Gemini-2.5-Flash, achieves an average score of only 65.08. Although instruction tuning boosts open-source performance: the best-performing open-source model, LLaMA-3.1-8B, shows significant room for improvement, highlighting the complexity of temporal analysis for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。