arXiv:2608.18918cs.LG2026-08

传统降维方法会遗漏弱耦合系统中的组件,新方法通过代数结构避免此问题。

Score the Algebra, Not the Span: Dimension Reduction for Transfer Operator Models of Dynamical Systems

  • 以生成的代数结构代替秩来评估降维,使乘积和幂运算无需额外成本。
  • 仅需10个代数坐标即可恢复被传统方法完全遗漏的组件,而传统方法在k<100时失效。
  • 适合建模弱耦合物理与生物系统,尤其适用于需预测隐藏组件的研究者。

对由多个弱相互作用组分构成的动力系统进行降维时,传统基于谱的方法(如主模式分析)可能需要指数级多的模式,或完全丢失某个组分——该组分在模型中消失而非被粗粒化表示,其任意函数都无法准确预测。这种现象称为线性遮蔽。原因是基于秩的模型每个模式需一个坐标。本文提出以坐标生成的σ-代数为评分标准,使乘积与幂运算免费,且组分成本仅取决于其生成器,而非所有相互作用。该准则为嵌入当前与未来之间的χ²散度,具有预算保证:两倍于动力系统内在维度的坐标即足以构建包含算子全部谱的代数嵌入,支持无限秩。变分形式可使用现成估计器,当判别器限制为双线性类时,退化为跨度上的VAMP评分,表明秩基方法属于同一家族的极端情况。在多个公开基准系统的组合上验证了该目标,示例显示:在所有秩k<100时,秩基方法完全忽略被遮蔽组分;而仅需10个代数坐标即可恢复全部组分。此外,所得代数表示能从少量标签预测被遮蔽组分,而直接从高维观测或从VAMP特征回归则失败。

原文摘要 · Abstract (English)

Dimension reduction for dynamical systems is standard practice, and the standard route is spectral: model the transfer (Koopman) operator by its leading modes. We show that on systems assembled from several weakly interacting components --- a structure common in physical and biological settings --- this may either require an exponential number of modes, or drop an entire component: the component is absent from the model rather than modeled coarsely, and no function of it can be predicted at any accuracy. We call this linear masking. The cause is that a rank-based model pays one coordinate per mode. We propose to score instead the $σ$-algebra the coordinates generate, so that products and powers come free and a component's cost is governed only by its generators rather than by all its interactions. The criterion is a $χ^2$-divergence between the embedded present and future, and it carries a budget guarantee: twice the intrinsic dimension of the dynamics is enough coordinates for an embedding whose algebra carries the operator's entire spectrum, with its full infinite rank. In variational form the criterion admits off-the-shelf estimators, and restricting its critic to the bilinear class returns the VAMP score on the span, so rank-based methods are one end of the same family. We demonstrate the proposed objective on a composite of published benchmark systems. We exhibit examples where the rank-based methods completely miss the masked components at all ranks $k<100$, while ten algebra coordinates recover all of them. In addition, the resulting algebra representation supports predicting the masked components from few labels, while direct regression from the high-dimensional observation or from the VAMP features fail.

动力系统降维代数结构谱方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。