用新方法评估生成的股市订单簿数据,更准捕捉复杂结构。
LOB-ID: Evaluating Synthetic Market Data by Inception Distances

- 基于深度嵌入构建评估框架,适配订单簿数据特性
- 在五种模型上验证,结果与结构保留程度一致
- 对数据扰动更敏感,适合检测生成质量缺陷
生成式模型在限价订单簿(LOB)数据方面进展迅速,但现有评估多依赖典型事实和特定市场统计量,难以捕捉订单簿轨迹的联合时间与跨层级结构。本文提出LOB-ID,一种基于嵌入的评估框架,将弗雷谢特起源距离(FID)与蒙日起源距离(MIND)适配至LOB数据。通过在五个股票四个月的二级订单簿数据上训练DeepLOB架构获取领域专用嵌入,验证了LOB-ID在时间、标的和嵌入检查点上的稳定性,并在受控扰动下呈现单调上升趋势。进一步设计了基于矩匹配的对抗攻击和深层订单簿扰动,以逃避统计评估,结果显示MIND对两类扰动仍保持显著敏感性。最后对五种生成模型(含随机基线与深度学习方法)进行评分,发现LOB-ID排序与各模型所捕捉的联合时空结构一致。
原文摘要 · Abstract (English)
Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected market statistics. These measures provide useful diagnostics but may not capture the joint temporal and cross-level structure of order-book trajectories. We introduce LOB-ID, an embedding-based framework that adapts the Fréchet Inception Distance (FID) and Monge Inception Distance (MIND) to LOB data. To obtain domain-specific embeddings, we train the DeepLOB architecture on four months of Level-2 order-book data for five equities. We show that LOB-ID is stable across time, instruments, and embedding checkpoints, and rises monotonically under controlled distortions. We then construct a moment-matching attack against FID and a deep-book perturbation that evades statistic-based evaluation. MIND remains substantially more sensitive to both distortions. Finally, we score five generative LOB models, spanning stochastic baselines and deep learning approaches, and find that LOB-ID ranks them in line with the joint temporal and cross-level structure each captures by construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。