针对时序数据优化分词的归一化方法,提升大模型预测精度
TOKON: TOKenization-Optimized Normalization for time series analysis with a large language model
- 基于分词特性设计归一化方法,将时序元素压缩为单个标记
- 多步预测RMSE降低7%至18%,具体效果依赖数据集与提示策略
- 适合使用大模型进行时序预测的研究者与工程应用
尽管大语言模型快速向通用人工智能演进,其在时序数据分析中的应用仍受限。为此,我们提出一种新型归一化技术——分词优化归一化(TOKON),该方法充分考虑分词的本质特性,将每个时序元素表示为单一标记,使令牌数量减少2到3倍。同时,我们设计了一种新的时序预测提示方法,称为带关怀的时序预测(TFSC),以进一步提升预测性能。实验结果表明,TOKON在多步预测中使均方根误差(RMSE)降低约7%至18%,具体取决于数据集和提示方式。此外,当与TOKON结合使用时,TFSC在某些数据集上表现出更优的预测准确率。
原文摘要 · Abstract (English)
While large language models have rapidly evolved towards general artificial intelligence, their versatility in analyzing time series data remains limited. To address this limitation, we propose a novel normalization technique that considers the inherent nature of tokenization. The proposed Tokenization-Optimized Normalization (TOKON) simplifies time series data by representing each element with a single token, effectively reducing the number of tokens by 2 to 3 times. Additionally, we introduce a novel prompt for time series forecasting, termed Time Series Forecasting with Care (TFSC), to further enhance forecasting performance. Experimental results demonstrate that TOKON improves root mean square error (RMSE) for multi-step forecasting by approximately 7% to 18%, depending on the dataset and prompting method. Furthermore, TFSC, when used in conjunction with TOKON, shows additional improvements in forecasting accuracy for certain datasets
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。