arXiv:2512.11251cs.LG2025-12被引 14

用10万条时间序列数据训练模型,让AI自动读懂并描述各类数据趋势。

Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language

  • 构建新型智能工作流,结合统计分析与大模型生成趋势描述
  • 在20个数据集上测试,表现优于LLaVA和GPT-4等主流模型
  • 适合需要快速洞察时间序列的科研与工业用户

时间序列数据在环境分析、农业、交通和金融等多个领域至关重要,但从中挖掘洞察通常依赖深厚的领域知识,过程耗时且人力密集。本文提出 extbf{Insight Miner},一个大规模多模态模型(LMM),旨在生成富含领域知识的高质量、全面的时间序列描述。为此,我们引入 extbf{TS-Insights} ootnote{可访问:https://huggingface.co/datasets/zhykoties/time-series-language-alignment},首个通用领域的时间序列与语言对齐数据集。该数据集包含从20个预测数据集中采样的10万条时间序列窗口。我们通过一种新颖的 extbf{代理式工作流} 构建该数据集:先使用统计工具提取原始时间序列特征,再用GPT-4将其合成连贯的趋势描述。在TS-Insights上进行指令微调后,Insight Miner在生成时间序列描述与洞察方面超越了当前最先进的多模态模型,如LLaVA和GPT-4。研究结果表明,利用多模态大模型进行时间序列分析具有广阔前景,为实现大模型将时间序列作为原生输入模态奠定了基础。

原文摘要 · Abstract (English)

Time-series data is critical across many scientific and industrial domains, including environmental analysis, agriculture, transportation, and finance. However, mining insights from this data typically requires deep domain expertise, a process that is both time-consuming and labor-intensive. In this paper, we propose \textbf{Insight Miner}, a large-scale multimodal model (LMM) designed to generate high-quality, comprehensive time-series descriptions enriched with domain-specific knowledge. To facilitate this, we introduce \textbf{TS-Insights}\footnote{Available at \href{https://huggingface.co/datasets/zhykoties/time-series-language-alignment}{https://huggingface.co/datasets/zhykoties/time-series-language-alignment}.}, the first general-domain dataset for time series and language alignment. TS-Insights contains 100k time-series windows sampled from 20 forecasting datasets. We construct this dataset using a novel \textbf{agentic workflow}, where we use statistical tools to extract features from raw time series before synthesizing them into coherent trend descriptions with GPT-4. Following instruction tuning on TS-Insights, Insight Miner outperforms state-of-the-art multimodal models, such as LLaVA \citep{liu2023llava} and GPT-4, in generating time-series descriptions and insights. Our findings suggest a promising direction for leveraging LMMs in time series analysis, and serve as a foundational step toward enabling LLMs to interpret time series as a native input modality.

时间序列多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。