arXiv:2607.10810cs.LGcs.AI2026-07

用多阶段生成结果融合,提升罕见风险的估计精度

Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models

论文配图:Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models
图 1 · 摘自论文原文
  • 训练中保存多个检查点,测试时融合生成样本以平滑尾部波动
  • 在固定采样量下,尾部风险估计误差显著低于单检查点方法
  • 适合金融风控等依赖极端事件预测的应用场景

深度生成模型常被用作数据稀缺下的决策模拟器,但在风险敏感应用中,其价值取决于对罕见负面情景的捕捉能力。标准生成目标侧重整体分布拟合,导致低概率尾部易受局部优化噪声影响,在有限模拟预算下尾部相关函数不稳定。本文提出时序样本集成(DSI),一种测试时推理框架,通过融合随机训练轨迹中多个检查点的生成样本,构建检查点混合分布,以平均各检查点的尾部波动,而非依赖单一脆弱终点。我们通过有限预算下的偏差-方差理论形式化该机制。实证表明,在多元合成过程和高频交易数据上,相较于单检查点基线,DSI 在固定模拟预算下显著降低尾部估计误差,优于标准扩散模型及最先进的尾部感知基线,且无需修改生成目标。

原文摘要 · Abstract (English)

Deep generative models are increasingly used as simulators for downstream decision-making under data scarcity, but in risk-sensitive applications their usefulness depends on rare adverse scenarios rather than typical samples. Standard generative objectives prioritize bulk distributional fidelity, leaving low-probability tails vulnerable to localized optimization noise and making tail-dependent functionals unstable under finite simulation budgets. We introduce Diachronic Sample Integration (DSI), a test-time inference framework that ensembles generated samples across checkpoints from a stochastic training trajectory. DSI targets a checkpoint-mixture distribution that averages checkpoint-specific tail fluctuations rather than relying on a single brittle endpoint. We formalize this mechanism through a finite-budget bias-variance theory. Empirically, across multivariate synthetic processes and high-frequency trading data, DSI substantially reduces tail-estimation error compared to single-checkpoint baselines under fixed simulation budgets, outperforming standard diffusion and state-of-the-art tail-aware baselines without modifying the generative objective.

生成模型风险估计尾部分析金融风控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。