arXiv:2606.13823cs.LGeess.SP2026-06

提出可提前判断时间序列嵌入是否有效的无训练准则

A Stationarity-and-Coupling Criterion for Training-Free Time-Lagged Spectral Embeddings of Multivariate Time Series

论文配图:A Stationarity-and-Coupling Criterion for Training-Free Time-Lagged Spectral Embeddings of Multivariate Time Series
图 1 · 摘自论文原文
  • 基于时滞相关矩阵与耦合性分析,构建无需训练的嵌入方法
  • 在4个满足条件的数据集上达到88.5%准确率,仅用单核CPU
  • 预检两步:平稳性与功率饱和,提前预测模型能否生效

我们研究多变量时间序列的无训练固定长度描述符,关注其何时具备有效性。核心是基于时滞相关矩阵截断至Marchenko-Pastur边缘(保留信号特征特征值),并通过通道间余弦相似度分类,零参数学习。在平稳高斯VAR(1)模型下,证明当信号近似平稳且类别信息存在于跨通道时序耦合而非单通道能量时,该描述符$D(τ)$可有效区分两类。推导出三类结论:可区分性条件、静态(τ=0)协方差退化为随机水平的原因,以及仅靠功率差异的平稳范式会破坏描述符性能。提出可操作的预飞行检测:增强Dickey-Fuller平稳性检验与功率基线饱和检查。在混合数据集验证中,四个满足条件的范式(Sleep-EDF、BCI-IV-2a、MIT-BIH、ESC-50)表现媲美强基线,20人留一主体测试达88.5±4.5%,仅需单核计算;三个违反条件的范式(非平稳ERP、金融波动、可穿戴压力)则如预测般失败,这些负例更具信息量。强调$D(τ)$并非最精确表示,其价值在于紧凑、无训练且适用范围可预先确定。

原文摘要 · Abstract (English)

We study training-free fixed-length descriptors for multivariate time series and ask not merely whether such a descriptor performs well, but when it can be expected to work at all. Our object of study is $D(τ)$, built from a time-lagged correlation matrix truncated at the Marchenko-Pastur edge so that only signal-bearing eigenvalues survive and classified by cosine similarity to class centroids with zero learned parameters. The central contribution is not the descriptor but a falsifiable applicability criterion for it. Working from a stationary Gaussian VAR(1) model, we argue that $D(τ)$ separates two classes when the signals are approximately stationary and the class information lives in their cross-channel temporal coupling rather than in marginal per-channel power. We derive, semi-formally, three consequences: a distinguishability condition, why the static ($τ=0$) covariance collapses to chance, and why a stationary but power-discriminated paradigm defeats the descriptor. The criterion is operational: a two-part pre-flight test -- an augmented Dickey-Fuller stationarity check and a power-baseline saturation check -- predicts applicability before any training. We validate both halves on a mixed assortment. On four paradigms that satisfy the criterion (Sleep-EDF, BCI-IV-2a, MIT-BIH, ESC-50) the descriptor is competitive with strong baselines at a fraction of their cost, reaching $88.5\pm4.5\%$ under 20-subject leave-one-subject-out on Sleep-EDF on a single CPU thread. On three that violate it -- non-stationary ERPs, and financial-volatility and wearable-stress regimes that are power-discriminated -- it fails exactly as the pre-flight predicts, and these negatives are the more informative half. We are explicit that $D(τ)$ is not the most accurate representation; its value is a compact, training-free embedding whose domain of validity is known in advance.

时间序列无训练嵌入判别准则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。