构建首个跨11领域的时序描述基准,评测大模型理解时间数据能力。
CaTS-Bench: Can Language Models Describe Time Series?
- 基于1746条人工重写的真实时序描述构建评估集
- 开源合成生成流水线,验证后使开源模型性能大幅提升
- 提供910道诊断题与专用度量,适合评估数值推理能力
时序描述任务要求将时间序列用自然语言表达,涉及数值推理、趋势解读和上下文理解。现有基准多依赖全合成或通用描述,常忽略元数据与视觉呈现。我们提出CaTS-Bench,一个覆盖11个不同领域的上下文感知时序推理综合性基准,核心为1746条高质量人工重写的黄金标准描述,用于衡量模型将数值趋势转化为可立即理解叙述的能力。为解决人工标注数据稀缺问题,我们还设计了一条可扩展的高保真合成描述生成流水线,并验证其质量。在该基准上评估主流视觉-语言模型,发现即使专有模型也难以捕捉时间描述中的数值细微差别;而对开源模型在合成数据上微调后,性能显著提升。最后,我们发布包含910道多选题的诊断套件,结合定制化数值指标,评估模型在时序领域特有的推理能力,确立了CaTS-Bench作为数值领域可信、基础的多模态文本生成评估框架。
原文摘要 · Abstract (English)
Time series captioning, the task of describing time series in natural language, requires numeric and temporal reasoning, trend interpretation, and contextual understanding. Existing benchmarks, however, often rely on fully synthetic or generic captions, and typically neglect metadata and visual representations. We introduce CaTS-Bench, a comprehensive benchmark for Context-aware Time Series reasoning across 11 diverse domains, centered on a gold-standard evaluation set of 1746 human-rewritten captions that measure how effectively models translate numeric trends into immediately interpretable narratives. To address the scarcity of human-annotated data, we also propose a scalable pipeline for generating high-fidelity synthetic captions, the quality of which we validate. We evaluate leading Vision-Language Models on our benchmark, revealing that even proprietary models struggle to capture numeric nuances in temporal descriptions, while finetuning open-source models on synthetic data yields substantial performance gains. Finally, we release a diagnostic suite of 910 multiple-choice questions and use tailored numeric metrics to gauge time-series-specific reasoning capabilities, establishing CaTS-Bench as a reliable foundation for grounded, multimodal text generation in numeric domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。