对比三种生成模型,找出最像真实金融数据的合成方法。
Synthetic Financial Data Generation for Enhanced Financial Modelling
- 用统一框架评估三种生成模型的逼真度和时序结构。
- TimeGAN在仿真度和时间一致性上最优,MMD低至1.84e-3。
- 适合需要高保真数据的金融建模与风险测试的人看。
金融领域因数据稀缺和保密性问题,常阻碍模型开发与稳健测试。本文提出一个统一的多标准评估框架,应用于三种代表性生成范式:统计类的ARIMA-GARCH基线、变分自编码器(VAEs)和时序生成对抗网络(TimeGAN)。基于标普500日度历史数据,评估其保真度(最大均值差异,MMD)、时间结构(自相关与波动聚集)以及下游任务实用性,包括均值-方差投资组合优化与波动率预测。实证结果表明:ARIMA-GARCH能捕捉线性趋势与条件异方差,但无法再现非线性动态;VAEs生成平滑轨迹,低估极端事件;TimeGAN在真实感与时序连贯性间取得最佳平衡(如平均5次种子下MMD达1.84e-3)。最后,根据应用需求与计算约束,给出模型选择建议。本文提出的统一评估协议与可复现代码库旨在推动合成金融数据研究的标准化。
原文摘要 · Abstract (English)
Data scarcity and confidentiality in finance often impede model development and robust testing. This paper presents a unified multi-criteria evaluation framework for synthetic financial data and applies it to three representative generative paradigms: the statistical ARIMA-GARCH baseline, Variational Autoencoders (VAEs), and Time-series Generative Adversarial Networks (TimeGAN). Using historical S and P 500 daily data, we evaluate fidelity (Maximum Mean Discrepancy, MMD), temporal structure (autocorrelation and volatility clustering), and practical utility in downstream tasks, specifically mean-variance portfolio optimization and volatility forecasting. Empirical results indicate that ARIMA-GARCH captures linear trends and conditional volatility but fails to reproduce nonlinear dynamics; VAEs produce smooth trajectories that underestimate extreme events; and TimeGAN achieves the best trade-off between realism and temporal coherence (e.g., TimeGAN attained the lowest MMD: 1.84e-3, average over 5 seeds). Finally, we articulate practical guidelines for selecting generative models according to application needs and computational constraints. Our unified evaluation protocol and reproducible codebase aim to standardize benchmarking in synthetic financial data research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。