用混沌时间序列生成数据,实现金融预测的零样本大模型训练。
Scaling Law for Large-Scale Pre-Training Using Chaotic Time Series and Predictability in Financial Time Series
- 通过合成混沌时序并重采样,构建金融数据训练集。
- 100亿样本预训练后,零样本预测在比特币上优于自相关模型。
- 发现预测性能随样本量指数增长的缩放规律,适合计算资源充足的研究者。
时间序列预测在气象、交通、电力、经济、金融等领域至关重要,尤其金融资产收益预测极具挑战性。已有研究提出适用于多种任务的时间序列基础模型。鉴于真实时间序列具有混沌特性,研究人员开发了人工生成合成混沌时间序列的方法,构建多样化数据集并用于模型训练。本文提出一种新方法:通过生成人工混沌时间序列,结合重采样技术模拟金融时序数据,并作为训练样本。扩大重采样间隔以延长预测时域,采用每种情形下100亿个训练样本进行大规模预训练。随后利用实际比特币交易数据构建多时间尺度测试集,对预训练模型进行零样本预测。基于预测结果评估简单交易策略的盈利能力,显示其显著优于自相关模型。在大规模预训练过程中,观察到类似缩放定律的现象:通过指数级增加训练样本数量,可在更长预测时域内达到特定预测性能水平。若该缩放定律在各类混沌模型中均成立,则表明可通过投入大量算力实现近未来事件的预测。未来研究应聚焦于更大规模训练,并验证该缩放定律在多样混沌模型中的适用性。
原文摘要 · Abstract (English)
Time series forecasting plays a critical role in decision-making processes across diverse fields including meteorology, traffic, electricity, economics, finance, and so on. Especially, predicting returns on financial instruments is a challenging problem. Some researchers have proposed time series foundation models applicable to various forecasting tasks. Simultaneously, based on the recognition that real-world time series exhibit chaotic properties, methods have been developed to artificially generate synthetic chaotic time series, construct diverse datasets and train models. In this study, we propose a methodology for modeling financial time series by generating artificial chaotic time series and applying resampling techniques to simulate financial time series data, which we then use as training samples. Increasing the resampling interval to extend predictive horizons, we conducted large-scale pre-training using 10 billion training samples for each case. We subsequently created test datasets for multiple timeframes using actual Bitcoin trade data and performed zero-shot prediction without re-training the pre-trained model. The results of evaluating the profitability of a simple trading strategy based on these predictions demonstrated significant performance improvements over autocorrelation models. During the large-scale pre-training process, we observed a scaling law-like phenomenon that we can achieve predictive performance at a certain level with extended predictive horizons for chaotic time series by increasing the number of training samples exponentially. If this scaling law proves robust and holds true across various chaotic models, it suggests the potential to predict near-future events by investing substantial computational resources. Future research should focus on further large-scale training and verifying the applicability of this scaling law to diverse chaotic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。