arXiv:2604.14498cs.AIcs.LG2026-04被引 1

合成数据能提升金融模型性能,但仅在方差主导场景有效。

Improving Machine Learning Performance with Synthetic Augmentation

论文配图:Improving Machine Learning Performance with Synthetic Augmentation
图 1 · 摘自论文原文
  • 将合成数据视为训练分布的修正,揭示其引入偏差-方差权衡
  • 仅在波动率预测等方差主导任务中提升性能,方向预测反而下降
  • 适用于关注高维金融建模与数据稀缺问题的研究者

合成数据在金融机器学习中被广泛用于缓解数据稀缺问题,但其统计作用仍不明确。本文将合成增强形式化为有效训练分布的修改,发现其引发结构化的偏差-方差权衡:额外样本虽可降低估计误差,但若合成分布偏离评估相关区域,可能改变总体目标函数。为分离信息增益与单纯样本量效应,提出大小匹配的零增强基线和有限样本、非参数的块置换检验,该方法在弱时间依赖下依然有效。我们在受控马尔可夫切换环境及真实金融数据集(包括高频期权交易数据和日度股票面板)上评估,涵盖自助法、基于拷贝的模型、变分自编码器、扩散模型和TimeGAN等多种生成器。实验考察了增强比例、模型容量、任务类型、状态稀有性及信噪比。结果表明,合成数据仅在方差主导的场景(如持续波动率预测)中有效,而在偏差主导的任务(如近似有效市场方向预测)中反而损害性能。针对稀有状态的靶向增强可提升特定指标,但可能违背无条件置换推断。研究提供了合成数据改善金融学习性能的结构性判断依据。

原文摘要 · Abstract (English)

Synthetic augmentation is increasingly used to mitigate data scarcity in financial machine learning, yet its statistical role remains poorly understood. We formalize synthetic augmentation as a modification of the effective training distribution and show that it induces a structural bias--variance trade-off: while additional samples may reduce estimation error, they may also shift the population objective whenever the synthetic distribution deviates from regions relevant under evaluation. To isolate informational gains from mechanical sample-size effects, we introduce a size-matched null augmentation and a finite-sample, non-parametric block permutation test that remains valid under weak temporal dependence. We evaluate this framework in both controlled Markov-switching environments and real financial datasets, including high-frequency option trade data and a daily equity panel. Across generators spanning bootstrap, copula-based models, variational autoencoders, diffusion models, and TimeGAN, we vary augmentation ratio, model capacity, task type, regime rarity, and signal-to-noise. We show that synthetic augmentation is beneficial only in variance-dominant regimes, such as persistent volatility forecasting-while it deteriorates performance in bias-dominant settings, including near-efficient directional prediction. Rare-regime targeting can improve domain-specific metrics but may conflict with unconditional permutation inference. Our results provide a structural perspective on when synthetic data improves financial learning performance and when it induces persistent distributional distortion.

合成数据金融建模偏差-方差增强方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。