用路径签名与注意力机制,高效估计分数阶布朗运动SDE的参数。
SigMA: Path Signatures and Multi-head Attention for Learning Parameters in fBm-driven SDEs
- 将路径签名与多头注意力结合,构建新型神经架构SigMA。
- 在合成数据和真实数据上均优于CNN、LSTM等基线模型。
- 适合处理具有长程依赖的金融与工程系统参数估计。
由分数阶布朗运动(fBm)驱动的随机微分方程(SDE)越来越多地用于建模具有粗糙动态和长程依赖的系统,如量化金融和可靠性工程。然而,这些过程是非马尔可夫且缺乏半鞅结构,导致许多经典参数估计方法无法应用或计算不可行。本文探讨两个核心问题:(i) 将路径签名整合到深度学习架构中能否改善估计精度与模型复杂度之间的权衡;(ii) 如何设计有效的架构以利用签名作为特征映射。我们提出SigMA(Signature Multi-head Attention),一种结合路径签名与多头自注意力的神经架构,包含卷积预处理层和多层感知机以实现有效特征编码。SigMA从分数阶布朗运动、分数阶奥恩斯坦-乌伦贝克过程及粗糙赫斯顿模型等生成的合成路径中学习模型参数,尤其关注赫斯特参数估计与多参数联合推断,并在未见轨迹上表现出强泛化能力。在合成数据及两个真实数据集(股票指数已实现波动率、锂离子电池退化)上的大量实验表明,SigMA在精度、鲁棒性和模型紧凑性方面持续优于CNN、LSTM、基础Transformer和Deep Signature基线。结果表明,将签名变换与基于注意力的架构结合,为具有粗糙或持久时间结构的随机系统提供了有效且可扩展的参数推断框架。
原文摘要 · Abstract (English)
Stochastic differential equations (SDEs) driven by fractional Brownian motion (fBm) are increasingly used to model systems with rough dynamics and long-range dependence, such as those arising in quantitative finance and reliability engineering. However, these processes are non-Markovian and lack a semimartingale structure, rendering many classical parameter estimation techniques inapplicable or computationally intractable beyond very specific cases. This work investigates two central questions: (i) whether integrating path signatures into deep learning architectures can improve the trade-off between estimation accuracy and model complexity, and (ii) what constitutes an effective architecture for leveraging signatures as feature maps. We introduce SigMA (Signature Multi-head Attention), a neural architecture that integrates path signatures with multi-head self-attention, supported by a convolutional preprocessing layer and a multilayer perceptron for effective feature encoding. SigMA learns model parameters from synthetically generated paths of fBm-driven SDEs, including fractional Brownian motion, fractional Ornstein-Uhlenbeck, and rough Heston models, with a particular focus on estimating the Hurst parameter and on joint multi-parameter inference, and it generalizes robustly to unseen trajectories. Extensive experiments on synthetic data and two real-world datasets (i.e., equity-index realized volatility and Li-ion battery degradation) show that SigMA consistently outperforms CNN, LSTM, vanilla Transformer, and Deep Signature baselines in accuracy, robustness, and model compactness. These results demonstrate that combining signature transforms with attention-based architectures provides an effective and scalable framework for parameter inference in stochastic systems with rough or persistent temporal structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。