Fiaingen生成的金融时间序列数据质量高,训练效果接近真实数据。
Fiaingen: A financial time series generative method matching real-world data quality
- 基于降维空间重叠度设计生成方法,提升合成数据逼真度。
- 合成数据使下游模型性能接近真实数据,生成耗时仅数秒。
- 适合金融量化研究、交易模型训练等需要高质量数据场景。
数据对推动金融领域机器学习研究与应用至关重要,但真实世界数据在数量、质量和多样性上仍受限,尤其各类金融资产的数据短缺直接影响机器学习模型在投资与交易中的表现。生成方法可缓解此问题。本文提出一套新型时间序列生成技术(命名为Fiaingen),并从三个维度评估其性能:(a) 实际与合成数据在降维空间中的重叠程度;(b) 在下游机器学习任务上的表现;(c) 运行效率。实验表明,Fiaingen在三项指标上均达到当前最优水平。生成的合成数据更贴近原始时间序列特征,且生成时间仅需数秒,具备良好可扩展性。使用该数据训练的模型性能接近以真实数据训练的模型。
原文摘要 · Abstract (English)
Data is vital in enabling machine learning models to advance research and practical applications in finance, where accurate and robust models are essential for investment and trading decision-making. However, real-world data is limited despite its quantity, quality, and variety. The data shortage of various financial assets directly hinders the performance of machine learning models designed to trade and invest in these assets. Generative methods can mitigate this shortage. In this paper, we introduce a set of novel techniques for time series data generation (we name them Fiaingen) and assess their performance across three criteria: (a) overlap of real-world and synthetic data on a reduced dimensionality space, (b) performance on downstream machine learning tasks, and (c) runtime performance. Our experiments demonstrate that the methods achieve state-of-the-art performance across the three criteria listed above. Synthetic data generated with Fiaingen methods more closely mirrors the original time series data while keeping data generation time close to seconds - ensuring the scalability of the proposed approach. Furthermore, models trained on it achieve performance close to those trained with real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。