用拓扑结构增强时间序列注意力,提升预测精度。
Global and Local Topology-Aware Attention with Persistent Homology and Euler Biases for Time-Series Forecasting

- 引入持久同调与欧拉偏置,显式建模时间序列的几何结构。
- 在真实数据上平均降低47.8%误差,显著优于基线模型。
- 适合关注时间序列内在几何规律的研究者或工业预测场景。
科学时间序列常包含连通性、环状结构、壳状几何、方向变化及非线性邻域等可预测的几何结构,而标准点积注意力无法显式表示。本文提出一种拓扑感知注意力框架,通过持久同调(H0-H2)、锚定欧拉特征变换和核-希尔伯特通道,在注意力得分中融入此类结构。验证门控的局部残差仅在保留验证数据支持时才引入局部拓扑信号,包括Zeng风格的局部H0分量。采用无泄露协议进行精确维托里斯-里普斯计算与平滑拓扑替代,训练仅校准,验证仅选择,测试仅报告。在三类架构(轻量注意力/岭回归、PatchTSTForRegression、TimeSeriesTransformerForPrediction)上评估,涵盖合成基准与真实数据(CO2、S&P 500回报窗口几何、NASA IMS轴承退化)。审计使用配对比较,含7个数据集单位、3个随机种子、3个时序划分,每架构63组配对,共189组。当几何结构具有预测性时,拓扑感知模型表现更优,结果异质性强:轻量注意力/岭回归在46/63单位改善,均方根误差降低12.5%,配对随机化p=7.2e-4;PatchTST在33单位改善,20单位持平,误差降低23.5%,p=3.5e-5;TimeSeriesTransformer在47单位改善,误差降低47.8%,p<1e-4。结果支持拓扑作为可验证、适配架构的归纳偏置。
原文摘要 · Abstract (English)
Scientific time series often encode predictive geometric structure, including connectivity, cycles, shell-like geometry, directional changes, and nonlinear neighborhoods, that standard dot-product attention does not explicitly represent. We introduce a topology-aware attention framework that adds such structure to attention logits using persistent homology (H0-H2), anchored Euler characteristic transforms, and kernel-Hilbert channels. A validation-gated local residual captures local topological signals, including a Zeng-style local H0 component, only when held-out validation data support the correction. Exact Vietoris-Rips computations and smooth topological surrogates are evaluated under a no-leakage protocol with train-only calibration, validation-only selection, and test-only reporting. We evaluate guarded topology-aware variants across three architecture families: lightweight attention/Ridge, PatchTSTForRegression, and TimeSeriesTransformerForPrediction. Experiments include synthetic benchmarks isolating higher-order topology and real datasets covering CO2, S&P 500 return-window geometry, and NASA IMS bearing degradation. The audit uses matched paired comparisons across seven dataset units, three random seeds, and three chronological splits, giving 63 paired units per architecture and 189 paired units overall. Topology-aware models show positive paired effects when geometry is predictive, with heterogeneous magnitude across datasets and architectures. Lightweight attention/Ridge improves in 46 of 63 units, with mean relative RMSE reduction of 12.5% and paired randomization p=7.2e-4; PatchTST improves in 33 units and retains the baseline in 20 units, with 23.5% reduction and p=3.5e-5; and TimeSeriesTransformer improves in 47 units, with 47.8% reduction and p<1e-4. The results support topology as a validation-selected, architecture-compatible inductive bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。