构建首个长时序多任务检索基准,测试模型对超长时间序列的推理能力。
TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning
- 设计跨10个领域的问答任务,时序跨度达100秒至24小时。
- 现有模型在长时序下准确率显著下降,部分无法处理超过100秒的数据。
- 基于代理的检索框架在9项任务上超越主流模型,适合长时序推理研究者。
时间序列语言模型(TSLMs)有望实现对真实世界时序数据的推理,但其在长时序数据上的检索与推理能力尚未得到充分检验。我们提出TS-Haystack,一个涵盖10个领域、包含10个事件锚定的问答任务的多任务检索基准,时序上下文长度从100秒到24小时不等,覆盖直接检索、时间推理、多步推理和上下文异常检测。现有TSLMs表现出严重的长上下文退化:准确率随上下文长度增加而下降;直接分词模型在高频信号下超过100秒即显存不足;时间间隔锚定任务在时序长度增加时准确率趋近于零,与文本和多模态长上下文检索的研究结果一致。采用专用时间序列分类工具的代理式检索框架,在9项任务上达到或超越当前最优TSLMs表现,表明代理式检索是长上下文TSLM的有前景方向。
原文摘要 · Abstract (English)
Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We introduce TS-Haystack, a multi-domain retrieval benchmark with ten event-grounded question-answering tasks over contexts from 100 seconds to 24 hours, spanning direct retrieval, temporal reasoning, multi-step reasoning, and contextual anomaly detection. Existing TSLMs exhibit severe long-context degradation: accuracy declines with context length, direct-tokenization models run out of memory beyond 100 seconds on high-rate signals, and time-interval-grounded tasks collapse toward near-zero accuracy when increasing the time-series lengths, aligning with existing literature on text and multi-modal long context retrieval. An agentic retrieval framework using specialized time-series classifier tools matches or outperforms SoTA TSLMs on 9 of 10 tasks, highlighting agentic retrieval as a promising approach for long-context TSLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。